Files
felhom.eu/STATUS.md
T
admin d30c2a51ed
gates / gates (push) Successful in 7s
R-224..R-228 CLOSED: registers, capability map, campaign annotation, STATUS
Five closed in controller v0.202.0 + agent v0.126.0, each with its live or
red-proof evidence in the row. Five explicitly still open and named as such
rather than left to inference: R-214, R-220, R-221, R-213, R-202 — and R-220 is
flagged as currently worked around BY HAND on the campaign venue, which is the
only reason an app could be deployed there.

The capability map's recovery row STAYS FAIL and says why: fixes are not a
re-walk, nothing walked a customer end to end, and the customer-facing messages
were NOT re-driven live because /recovery correctly retires itself once the old
data is set aside — restoring that state is the reconfiguration the task forbade.

The campaign document is ANNOTATED, not rewritten: it records what was true when
it ran, and that is its value.

workspace-CLAUDE.md gains comment-vs-code entry 9 — the escrow header said the
errors were 'DISTINCT on purpose' and named THREE situations while a fourth was
folded into one of them, and a green test named the defect and did not prevent
it because it asserted a STRING one layer below the merge.

ROADMAP needed no collapse — it carries no rows for these IDs.
2026-08-06 08:34:04 +02:00

187 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# STATUS — what works, what's broken, what's next
**Updated 2026-08-06.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
> items rather than extend it.
## What works right now
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
home network, and open apps from a launcher or a shared link. Backups run on their own to three
places — the machine's drive, a second drive, and an encrypted off-site copy.
**The backup promise is proved. The recovery JOURNEY is not.** On 5 August we built a brand-new
machine from the published disc, gave it three marked files, destroyed it, and tried to get them back
**the way a household would** — no shortcuts, no command line. The files came back **byte for byte
identical**, all three, including one with Hungarian accents in its name. But **the journey needed us
four times**, and the very first thing the machine did was tell the customer their correct recovery
code was wrong. *(CAMPAIGN 11)*
## What's broken
- **A machine installed today still gets the older in-house service, so it cannot open a recovery
package until you approve the newer one.** It is no longer *lied to* — it says plainly that the
machine cannot do this yet — but **approving the new service is one click from you**, and until then
such a machine also gets the cautious "we do not know why" wording rather than the helpful one.
*(R-216, R-223, R-224)*
- **Three things a rebuilt machine still cannot do by itself.** Its owner cannot re-attach their own
drives, so no app can be put back on its data; it cannot create a new recovery code at all; and the
screen at the machine itself never stops showing a stale pairing code. Each is understood, measured
and written down — none is fixed yet. *(R-220, R-221, R-214)*
- **Rebuilding a machine still throws away its off-site backup HISTORY.** The machine invents the key
that encrypts its own off-site backups, and a rebuilt machine invents a new one. **The good news:
the old key really is kept now — we proved it on a real machine today, for the first time**, and a
changed key raises an alarm the same day. *(R-193, R-198)*
- **The kept older backups cannot be opened — by anyone.** We keep the previous sealed package and
there is no way to open it. The screens now say exactly that and stop. **One place still promises
otherwise**: the older-backups card says they "may be restorable later with the matching code",
which is not true today. *(R-222, R-202)*
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
stopgap. *(R-95, R-87)*
## What last night's stress test found — and what we fixed this morning
We spent the night trying to break the recovery journey, then left the machine alone and watched it
run. **Nothing we did lost a byte.** When the customer chose "I do not want the old data", the old
backups were **set aside and not deleted** — we checked the far end of the wire and the 12.5 MB was
still there, to the byte. A wrong code was refused three times with nothing written and no lockout.
The machine's alarm fired when we switched it off and cleared itself when it came back. Overnight it
ran a full cycle on its own and made a fresh off-site copy without being asked.
**What it found: the machine still blamed the customer for failures that were not theirs.** Pull the
plug on our own central system and the customer was told their recovery code was bad — in three
hundredths of a second, when actually checking a code takes about one. The machine had not even
tried. **All of that is fixed and deployed** *(R-224, R-226, R-225, R-227, R-228)*:
- **When something on our side is down, we say so** — and we say plainly that the code was **not**
used, so it is still good. Proven on the real machine: with our hub unreachable the answer changed
from "your code is wrong" to "we could not reach the central system".
- **A customer who mistypes is told to check their typing again.** That message had become
unreachable on any machine that had been given a new code — exactly the machine that just recovered.
- **When we do not know why something failed, we say that**, and never guess the customer.
- **"0 snapshots · 0 GB" is gone** where the truth is "we have not read it yet".
- **The set-aside backups are visible again** — the machine says they are kept and not deleted, and
does **not** pretend they can be reopened, because today they cannot be.
**Still open, and worth knowing:** a rebuilt machine still cannot re-attach its own drives without us
*(R-220 — we are working around it by hand on the test machine right now)*, cannot create a new
recovery code *(R-221)*, and the screen at the machine still shows a stale pairing code *(R-214)*.
**The recovery journey is still recorded as FAILED** — these are fixes, not a re-walk, and it stays
failed until someone walks it end to end with no help from us.
## What shipped recently
- **2026-08-05 (late)** — **New machines now get current software again — and the disc image had been
quietly un-updatable for days.** Approving the newer in-house service turned out to be impossible on
its own: the system correctly refuses to publish a set where the pre-built machine image is older
than the software the fleet already runs, and that image had been behind since late July. **So the
image was rebuilt and both were published together.** A machine installed from now on lands on
current software and can open a recovery package on day one. *(R-223)*
- **2026-08-05 (evening)** — **Four sentences where there was one, and none of them blames you.**
"We did not accept your recovery code" used to appear when the code was wrong, when the machine
could not ask, when the store could not be read, and when the customer held the code for an older
backup we still keep. Only the first is the customer's doing. Also today: succeeding at recovery no
longer switches off the machine's own request for the thing it still needs; the screen now finishes
the job and shows what is in the backups instead of promising a list it could never produce; and the
recovery page can no longer be reached on a machine that never had backups.
*(R-216, R-217, R-218, R-219, R-222, R-215)*
- **2026-08-05** — **The recovery screen: a customer whose machine was rebuilt is now told, and shown
how.** Until today they had everything needed to get their data back and no way to find out — the
only route was a command line. The screen unlocks the backups and lists what is in them; it does
**not** restore anything, because unlocking and restoring are two different decisions and mixing
them would turn one clear moment into a wizard. Three ways out, none of them a dismiss button — and
the „most nem" option keeps the route to the data permanently visible in the backups area, because
a notice someone clicks past once is a notice that never happened. *(R-193)*
- **2026-08-05** — **The orphaned backups are deleted — and the list you were given was wrong, which
is why you were asked again.** You had approved "about 1.2 GB in two set-aside stores". Measured
before touching anything: there were **three** set-aside stores totalling **~1.45 GB** — and the
thing that was exactly 1.2 GB was demo-felhom's **live** store. Matching on the size would have
deleted a working backup. With the corrected list confirmed, all three were removed and both live
stores left alone; a real off-site backup ran successfully straight afterwards to prove nothing
working had been caught. *(R-212)*
- **2026-08-05** — **A rebuilt machine now asks for its storage credential, and the hub gives it back.**
The machine says plainly what it needs — it can tell it has been rebuilt, because its data area is
empty *and* the hub is holding a sealed recovery package for it — instead of leaving the hub to guess
from a silence that has four possible meanings. The hub waits long enough to be sure it is not a
restart, then **re-uses the credential it already holds** before creating a new one at the storage
provider. **The recovery ceremony stays manual, deliberately:** a credential can be replaced, your
recovery code cannot. *(R-204 item 4, R-193)*
- **2026-08-05** — **The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer
is unstuck.** The cause was leftover working data from building our own software — 157 GB of it,
growing about 5 GB a day, which nothing was allowed to delete. **148 GB came back in 86 seconds.**
The storage layer had already stopped accepting new copies of any volume onto that disk; that is
fixed the same day. A **30 GB ceiling** is now in place and was **proved to work by deliberately
overfilling it and watching it evict** — not by assuming the setting took. Two things that failed
quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved),
and the weekly cleanup is still forbidden from touching the thing that grows (next session).
*(R-205 … R-211)*
- **2026-08-05** — **The build data now lives on the second SSD, moved with nothing lost.** All 345
images, every saved volume and both development databases came through identical — checked before
the original was touched and again afterwards, and confirmed by running a real build on the moved
copy. **The server's own services never went down:** Gitea, the registry, the hub and the backup
system run on a separate system and stayed up throughout. The second SSD now also **reserves 80 GB**
for this, so the storage layer can no longer quietly claim the space and repeat what happened to the
first disk. The old copy is kept as the way back until the machine next restarts. *(R-209)*
- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time;
a credential re-issue no longer blocks off-site backups on a healthy machine; and the default
restore no longer quietly returns the wrong thing. *(R-204 items 13, R-196)*
- **2026-08-04 (night)** — **The drill PASSED**, end to end, on real hardware. *(R-201)*
- **2026-08-04** — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup
look in the same directory, and a backup that misses a folder marked essential reports *incomplete*
instead of success. *(R-203)*
- **2026-08-04** — The hub now keeps the off-site backup key when a machine re-seals, instead of only
the whole-machine one, and a machine can fetch its own sealed package back. *(R-198, R-199)*
- **2026-08-04** — The daily false alarm about David is gone; a vanished permission now repairs itself
**and says it had to**; the weekly off-site backup stopped reporting failure after a successful
upload. *(R-195, R-190, R-191)*
## What we're working on
- **Next: the retention proof.** The hub keeping the old sealed key when a machine re-seals is the one
remaining link that has **never run outside a test**. Proving it needs a second deliberate wipe on
the spare demo machine, and it is its own procedure. *(R-198)*
- **Then:** the orphaned-backup deletion you asked for (below), and the off-site copy the machine can
still erase. *(R-193, R-95)*
## Waiting on you
- **One thing to read after the machine next restarts — and nothing to do until then.** You told me
not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been
through a restart. It works right now and nothing was lost, but a restart is the one test that
matters for this kind of change, and it has not happened. **I made it check itself:** whenever the
machine next starts, for any reason, it writes a plain PASS or FAIL line to
`/var/log/felhom-store-postboot-check.log`. **If it says PASS, the old copy can be deleted and 34 GB
comes back.** Until then I have deliberately kept that old copy, which is the only quick way back —
it is why the disk sits at 54% rather than lower. *(R-209a)*
- **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and
62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old
(controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing
was deleted.** Worth knowing before you spend time on it: they only account for about 27 GB against
199 GB now free, so this is about clutter, not space. *(R-210)*
- **Nothing.**
- **The recovery screen you described has been priced, and it can be built.** A freshly installed
machine that finds a sealed package waiting should say so, offer a box for the recovery code, and
show what would come back before doing anything. One thing to weigh, deliberately not decided: that
screen is reachable by anyone with the household's dashboard password, and the preview reveals
backup dates and app names. *(R-193)*
- **The orphaned backups on the storage box — you said delete, and it is still owed.** About 1.2 GB
across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its
own session rather than riding along with other work. *(R-193)*
- *(decided 4 Aug)* You chose **not** to keep a copy of the backup key on the Proxmox host, which
makes the customer's own recovery code the only route back from a rebuild. *(R-193)*
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
session log; nothing suggests anyone else saw it. *(R-132)*
- **One small question, not urgent.** The automatic version check cannot see which version you have
told machines to install, only which ones exist. *(R-184)*