STATUS.md rebuilt (258 -> 85); R-245 re-filed as decided; R-246/247/248 filed

STATUS.md REBUILT FROM THE REGISTER, not trimmed. Its own header says one
screen; it had reached 258 lines, having been 83 four days ago.

The three named defects, all fixed:
  1. the "waiting on you" list asked the operator to decide the RECOVERY
     SCREEN, built and shipped 2026-08-05, and to approve an orphaned-backup
     deletion the register records as DONE the same day;
  2. a stray line reading only "- **Nothing.**" sat mid-list;
  3. the DooPlex infrastructure work was mixed in with the product's.

Infrastructure is now under ITS OWN HEADING rather than dropped, and the
reason is stated on the page: these are real asks that need the operator, but
they concern the machine this is built on, not what a customer receives.
Dropping them would lose real work; mixing them is why the page stopped being
readable.

The 100-line "what shipped recently" log is gone. That is what the per-repo
CHANGELOGs and the register are for, and restating it here is what made the
page grow back.

R-245 RE-FILED as a decision taken, not a question pending. It sat as
WAITING-ON-OPERATOR for a day with nothing actually pending - it was settled
on 2026-08-07. It keeps the whole reasoning and now carries the condition that
would REOPEN it, which the reasoning already named: QUOTA, old set-aside
history blocking new backups. A condition, not a calendar.

AUDIT OF EVERY WAITING-ON-OPERATOR ROW, parsing the state column exactly
rather than grepping for the phrase (which over-matches rows that merely
mention it): exactly ONE row carried it - R-245 - and it was a settled
decision. So zero rows were genuinely waiting, and the drift was caught while
it was still a single row.

R-246/R-247/R-248 file the read-only stale-blob spike's findings. R-242
updated: it recurred within a day, and shape (b) is now built - with the vouch
half explicitly still open on that row rather than being papered over.
This commit is contained in:
2026-08-07 12:59:52 +02:00
parent 3ca9a7bbe6
commit ceac5e0deb
2 changed files with 73 additions and 242 deletions
+67 -240
View File
@@ -1,258 +1,85 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-07.**
**Updated 2026-08-08.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
> items rather than extend it.
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical
> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something
> belongs in the register instead.
>
> *Rebuilt from the register on 2026-08-08. It had reached 258 lines; its "waiting on you" list asked
> for two things already shipped and carried a stray line reading only "Nothing."; and it mixed the
> DooPlex infrastructure work in with the product. The old "what shipped recently" log — 100 lines of
> it — is what the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.*
## What works right now
## What works
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
home network, and open apps from a launcher or a shared link. Backups run on their own to three
places — the machine's drive, a second drive, and an encrypted off-site copy.
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
sets their own password. They install apps from a catalogue of fifty-three, share files over the home
network, and open apps from a launcher or a shared link. Backups run on their own to three places — the
machine's drive, a second drive, and an encrypted off-site copy.
**The backup promise is proved again. The recovery JOURNEY still is not — but the fault that broke it
last night is now fixed.** On 67 August we built a brand-new machine from the published disc, gave it
three marked files, soaked it through a full night, destroyed it and rebuilt it. **The files came back
byte for byte identical**, all three, including one with Hungarian accents in its name. **But the
household had no way to ask for them:** the screen that takes their recovery code had switched itself
off, and the backup page offered to make them a *new* code instead.
**7 August — we asked which of two things was wrong, and the answer reversed the repair.** The obvious
reading was that the screen's rule was too narrow. It was not. **The screen was telling the truth**
there really was nothing openable with the key the machine held, **because the machine had made that
key itself**, on top of the sealed package we were already keeping for it. And it knew: it wrote *"the
sealed package does not cover the current key"* into its own log **thirty-five minutes before the
household looked**, then threw the answer away. Mending the screen would have hidden a machine quietly
making its own backups unopenable. *(R-241)*
**Fixed the same day.** The machine **stops making its own key** while we hold a package for it — the
repair that prevents the situation rather than tidying up after it. **The comparison it was already
making now decides whether to offer help**, instead of two indirect guesses that have each been wrong
in opposite directions. And **giving up the old backups became a finishable thing**: a **14-day
countdown** you can see and change your mind about, at the end of which the old backups *and* their
sealed package go together — so the question stops coming back because there is nothing left to ask
about, not because something is suppressing it. On the way past, the *"create a new recovery code"*
button is now **unavailable** while a recovery is outstanding, and the confirmation says plainly that
the old backups are deleted **on a date** rather than merely set aside.
**Two of our own mistakes were caught by tests rather than by reading the code**, which is the point of
having them: one would have warned a household that had switched off-site backups off, and one would
have brought back the very fault we were fixing.
**Still open, on purpose:** a machine held waiting for its code raises no alarm to *us* *(R-243)*;
nothing enforces that a release reaches a new machine *(R-242)*; and we deliberately did **not** build
automatic abandonment after 30 days *(R-245 — the reasoning is written down)*.
**Today's machines now get last week's fixes — approved 7 August.** *(R-239 — closed.)*
**The backup promise is proved.** A machine has been destroyed on purpose and its files came back byte
for byte identical — three separate times, including a filename with Hungarian accents.
## What's broken
- **A backup could report „✓ Rendben" while quietly leaving out an app the customer had just
chosen — FIXED 6 August.** Two things were wrong and only one had been guessed at. The machine's
own rule said *"a warning beside a success is read as a success"*, and it applied it to an app
missing a *folder* but not to an app left out *entirely*; that now counts too. And the real cause of
the case we saw: pressing „Távoli mentés most" while a backup was already running answered
„elindult" and then showed the **previous** run's green tick — so the customer read it as covering
their new app. It did not, and the restore refused minutes later. The button now says plainly that
it did not start anything.
- **A household still cannot get their own data back unaided.** Every individual link now works; no
single walk has completed end to end without someone stepping in. *(R-201)*
- **The machine's own screen keeps telling an already-paired box to pair itself** — 25 minutes after it
was paired, on a screen that promises it refreshes itself. *(R-214, R-235)*
- **A rebuilt machine cannot create a new recovery code at all.** *(R-221)*
- **A backup that covered nothing still calls itself „Sikeres".** The state is honest; the word is not.
*(R-240)*
- **A machine waiting for its recovery code raises no alarm to us.** It quietly stops making off-site
backups, and three separate safety nets each correctly decide it is not their business. The
household can see it; we cannot. *(R-243)*
- **The card offering to reopen set-aside backups promises more than we can deliver** — we keep the old
sealed package, but nothing can open it. *(R-202)*
- **Deleting a customer leaves rows behind** on every test machine ever torn down, while reporting a
clean teardown. No secrets involved, but it accumulates with each walk. *(R-244)*
- **Putting restored files back where they belong is still a manual step.** *(R-213)*
- **~~A machine installed today still gets the older in-house service~~ — RESOLVED 7 August.** The
newer service is packaged and approved, so a newly installed machine now gets it without anyone
touching the machine. *(R-216, R-223, R-224, R-239)*
- **Three things a rebuilt machine still cannot do by itself.** Its owner cannot re-attach their own
drives, so no app can be put back on its data; it cannot create a new recovery code at all; and the
screen at the machine itself never stops showing a stale pairing code. Each is understood, measured
and written down — none is fixed yet. *(R-220, R-221, R-214)*
- **~~Rebuilding a machine throws away its off-site backup history~~ — the CAUSE is fixed (7 August).**
A rebuilt machine used to invent a new encryption key over the top of the sealed package we hold for
it. It no longer does: while we hold a package, it waits for the household's recovery code instead.
The old key is kept and a changed key still raises an alarm the same day. *(R-193, R-198, R-241)*
- **The kept older backups cannot be opened — by anyone.** We keep the previous sealed package and
there is no way to open it. The screens now say exactly that and stop. **One place still promises
otherwise**: the older-backups card says they "may be restorable later with the matching code",
which is not true today. *(R-222, R-202)*
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
stopgap. *(R-95, R-87)*
## Found today
## Can a household get their data back on their own? Asked again today — still no, but nearer
We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it —
guest and both drives, as a hardware loss would — and tried to get them back the way a household
would. *(R-201, the re-walk)*
**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
backup, in **23 seconds**, through the customer's own restore screen.
**And much of the journey now works.** The machine showed the recovery screen **without being asked**,
told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
code, and accepted the real code first time. The emailed claim code worked first try.
**But it still needed us twice**, and a household has neither hand:
- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
only half of it was)*
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
back, so the restore screen stays empty. *(R-220)*
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
built and published but not approved for new machines**. We installed them by hand for this test. So
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
## What we fixed this morning, and what it did not fix
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
cleared itself. What it found was that **the machine blamed the customer for failures that were not
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
remaining dead ends are different ones.
## What shipped recently
- **2026-08-06 (later)** — **The page that says which machine may be wrecked had one wrong sentence,
and it has been checked against the machine rather than corrected from memory.** It claimed the
spare HP box keeps no off-site copy, and used that as the reason to aim risky backup tests at the
other one. The HP box **does** keep an off-site copy — two of them, the newest from 4 August, in
its own space on the off-site server. The sentence was true when it was written and went out of
date on 23 July. It matters because a belief that a machine holds nothing is how a destructive test
lands on one that does. Also today: the notes the assistant keeps had three statements that were
simply untrue (including that the demo boxes were still away) — those three are fixed, the rest are
now flagged automatically rather than rewritten by hand. *(R-229, R-230)*
- **2026-08-06** — **The assistant's own working notes were on one machine with no copy anywhere.**
Everything the assistant has learned about this system over months — where things live, which traps
cost us an incident — sat in a folder on the build machine that no backup touched and no repository
held. It is now included in that machine's nightly backup. **Two things you should know before
treating that as solved:** the copy lands on the **same physical disk** as the original, so it
survives a mistake but not a dead drive, and the build machine's backups have **no off-site copy at
all**. **A full survey of that machine's backup, done the same day, confirmed both and found two
more things worth knowing.** The good news first: it has run every night without missing a set, and
we pulled a file back out of it and checked it matched the original exactly — the first time that
has ever been demonstrated. The rest: **if a backup ever fails, nobody is told** — the alert was
configured but never given anywhere to send to — and **every copy it makes stays inside that one
box**, so it survives any single disk dying but not the room. Nothing was changed; the survey was
read-only and the decisions are yours. Also tidied the same day: a quarter of those notes had become unreachable — filed but listed
nowhere, so nothing would ever read them — and the instructions the assistant reads at the start of
every session were cut roughly in half, with a check added so they cannot quietly grow back. Nothing
was deleted. *(R-229, R-230, R-231)*
- **2026-08-05 (late)** — **New machines now get current software again — and the disc image had been
quietly un-updatable for days.** Approving the newer in-house service turned out to be impossible on
its own: the system correctly refuses to publish a set where the pre-built machine image is older
than the software the fleet already runs, and that image had been behind since late July. **So the
image was rebuilt and both were published together.** A machine installed from now on lands on
current software and can open a recovery package on day one. *(R-223)*
- **2026-08-05 (evening)** — **Four sentences where there was one, and none of them blames you.**
"We did not accept your recovery code" used to appear when the code was wrong, when the machine
could not ask, when the store could not be read, and when the customer held the code for an older
backup we still keep. Only the first is the customer's doing. Also today: succeeding at recovery no
longer switches off the machine's own request for the thing it still needs; the screen now finishes
the job and shows what is in the backups instead of promising a list it could never produce; and the
recovery page can no longer be reached on a machine that never had backups.
*(R-216, R-217, R-218, R-219, R-222, R-215)*
- **2026-08-05** — **The recovery screen: a customer whose machine was rebuilt is now told, and shown
how.** Until today they had everything needed to get their data back and no way to find out — the
only route was a command line. The screen unlocks the backups and lists what is in them; it does
**not** restore anything, because unlocking and restoring are two different decisions and mixing
them would turn one clear moment into a wizard. Three ways out, none of them a dismiss button — and
the „most nem" option keeps the route to the data permanently visible in the backups area, because
a notice someone clicks past once is a notice that never happened. *(R-193)*
- **2026-08-05** — **The orphaned backups are deleted — and the list you were given was wrong, which
is why you were asked again.** You had approved "about 1.2 GB in two set-aside stores". Measured
before touching anything: there were **three** set-aside stores totalling **~1.45 GB** — and the
thing that was exactly 1.2 GB was demo-felhom's **live** store. Matching on the size would have
deleted a working backup. With the corrected list confirmed, all three were removed and both live
stores left alone; a real off-site backup ran successfully straight afterwards to prove nothing
working had been caught. *(R-212)*
- **2026-08-05** — **A rebuilt machine now asks for its storage credential, and the hub gives it back.**
The machine says plainly what it needs — it can tell it has been rebuilt, because its data area is
empty *and* the hub is holding a sealed recovery package for it — instead of leaving the hub to guess
from a silence that has four possible meanings. The hub waits long enough to be sure it is not a
restart, then **re-uses the credential it already holds** before creating a new one at the storage
provider. **The recovery ceremony stays manual, deliberately:** a credential can be replaced, your
recovery code cannot. *(R-204 item 4, R-193)*
- **2026-08-05** — **The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer
is unstuck.** The cause was leftover working data from building our own software — 157 GB of it,
growing about 5 GB a day, which nothing was allowed to delete. **148 GB came back in 86 seconds.**
The storage layer had already stopped accepting new copies of any volume onto that disk; that is
fixed the same day. A **30 GB ceiling** is now in place and was **proved to work by deliberately
overfilling it and watching it evict** — not by assuming the setting took. Two things that failed
quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved),
and the weekly cleanup is still forbidden from touching the thing that grows (next session).
*(R-205 … R-211)*
- **2026-08-05** — **The build data now lives on the second SSD, moved with nothing lost.** All 345
images, every saved volume and both development databases came through identical — checked before
the original was touched and again afterwards, and confirmed by running a real build on the moved
copy. **The server's own services never went down:** Gitea, the registry, the hub and the backup
system run on a separate system and stayed up throughout. The second SSD now also **reserves 80 GB**
for this, so the storage layer can no longer quietly claim the space and repeat what happened to the
first disk. The old copy is kept as the way back until the machine next restarts. *(R-209)*
- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time;
a credential re-issue no longer blocks off-site backups on a healthy machine; and the default
restore no longer quietly returns the wrong thing. *(R-204 items 13, R-196)*
- **2026-08-04 (night)** — **The drill PASSED**, end to end, on real hardware. *(R-201)*
- **2026-08-04** — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup
look in the same directory, and a backup that misses a folder marked essential reports *incomplete*
instead of success. *(R-203)*
- **2026-08-04** — The hub now keeps the off-site backup key when a machine re-seals, instead of only
the whole-machine one, and a machine can fetch its own sealed package back. *(R-198, R-199)*
- **2026-08-04** — The daily false alarm about David is gone; a vanished permission now repairs itself
**and says it had to**; the weekly off-site backup stopped reporting failure after a successful
upload. *(R-195, R-190, R-191)*
- **A leftover flag has been silently switching off the new recovery detection on one demo machine
since 4 August.** A Re-issue during the recovery drill set it, using code we removed the next day.
**The flag is wrong** the sealed package does cover the key that machine is using, and the two
fingerprints match exactly. Because of it the hub withholds a figure the machine needs, so the
machine tells its owner *"create a new recovery code"* — the one act that would put their old backups
beyond reach. **A freshly installed machine cannot reach this state**, because nothing has set that
flag since 5 August. *(R-246, R-247, R-248)*
## What we're working on
- **Next: the retention proof.** The hub keeping the old sealed key when a machine re-seals is the one
remaining link that has **never run outside a test**. Proving it needs a second deliberate wipe on
the spare demo machine, and it is its own procedure. *(R-198)*
- **Then:** the orphaned-backup deletion you asked for (below), and the off-site copy the machine can
still erase. *(R-193, R-95)*
- **The walk that either finishes the arc or says why not.** It needs the base image current (done
today) and a machine whose recovery detection is not silently switched off (answered today). *(R-201)*
- **Proving the hub really keeps the old sealed key** when a machine re-seals. Never run outside a
test; needs a second deliberate wipe and its own session. *(R-198)*
## Waiting on you
- **Approve the new base image**, so every future installation carries this week's fixes. *(R-239,
R-242 — presented separately in this session.)*
- **One thing to read after the machine next restarts — and nothing to do until then.** You told me
not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been
through a restart. It works right now and nothing was lost, but a restart is the one test that
matters for this kind of change, and it has not happened. **I made it check itself:** whenever the
machine next starts, for any reason, it writes a plain PASS or FAIL line to
`/var/log/felhom-store-postboot-check.log`. **If it says PASS, the old copy can be deleted and 34 GB
comes back.** Until then I have deliberately kept that old copy, which is the only quick way back —
it is why the disk sits at 54% rather than lower. *(R-209a)*
- **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and
62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old
(controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing
was deleted.** Worth knowing before you spend time on it: they only account for about 27 GB against
199 GB now free, so this is about clutter, not space. *(R-210)*
- **Nothing.**
- **The recovery screen you described has been priced, and it can be built.** A freshly installed
machine that finds a sealed package waiting should say so, offer a box for the recovery code, and
show what would come back before doing anything. One thing to weigh, deliberately not decided: that
screen is reachable by anyone with the household's dashboard password, and the preview reveals
backup dates and app names. *(R-193)*
- **The orphaned backups on the storage box — you said delete, and it is still owed.** About 1.2 GB
across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its
own session rather than riding along with other work. *(R-193)*
- *(decided 4 Aug)* You chose **not** to keep a copy of the backup key on the Proxmox host, which
makes the customer's own recovery code the only route back from a rebuild. *(R-193)*
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
session log; nothing suggests anyone else saw it. *(R-132)*
- **One small question, not urgent.** The automatic version check cannot see which version you have
told machines to install, only which ones exist. *(R-184)*
*Nothing else is pending. R-245 — whether an undecided household is auto-abandoned after 30 days — was
**settled on 7 August** (we do not build it, and the reasoning is recorded). It has been re-filed as a
decision taken rather than a question sitting in your queue.*
## DooPlex infrastructure — separate from the product
*Kept under its own heading rather than dropped: these are real asks that need you, but they concern
the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped
being readable.*
- **DooPlex's own backup makes every copy inside the same box, and nothing says when it fails.**
*(R-232)*
- **193 old images exist only on this machine** and cannot be re-downloaded. About 27 GB against 199 GB
free — clutter, not space. Nothing deleted. *(R-210)*
- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests
anyone else saw it. *(R-132)*
- **One thing to read after DooPlex next restarts.** The move to the second SSD has never been through
a reboot; it now writes PASS or FAIL to `/var/log/felhom-store-postboot-check.log` on every start.
On PASS, 34 GB comes back. *(R-209a)*
- **Backup scripts on DooPlex are unversioned host state.** *(R-231)*
- **Instruction-file follow-ups**, each needing a decision rather than an edit. *(R-229, R-230)*
File diff suppressed because one or more lines are too long