feed748325
gates / gates (push) Successful in 18s
R-234 was filed as "toggling an app on leaves it without a bundle, so the first run skips it". Measured on demo-hp: that state does not survive a run — the off-site run's own pre-dump phase calls captureAllRecoveryUnits for every DEPLOYED stack, through admitApp, before the push, and a unit moved aside was RECREATED. The actual cause was the single-flight: the manual run was dropped because an earlier one was still going, runOffboxBackup returned nil, the handler had already answered "A tavoli mentes elindult", and the card then showed the PREVIOUS run's green verdict. Fixed in controller v0.205.0 and proven live on demo-hp: a second request while one is in flight now says "Mar fut egy tavoli mentes — ez a keres nem inditott ujat. A most lathato eredmeny meg a korabbi futase", as a flash_error. Independently, and a real gap on its own: a run that skipped an app the customer selected is now `incomplete`, not `ok`. Selected+deployed with no unit counts; selected-but-undeployed is named with what to do but does NOT count, because a box left amber by an app somebody removed is a status nobody reads. R-218's state field read REOPENED while the same row's body already recorded the fix shipped in v0.203.0 and proven live. Corrected to CLOSED, keeping the over-claim history — it is why the row is worded as it is. Capability map: the off-site capture row's `incomplete` sentence widened to cover a whole-app skip, and it still does not claim a newly-selected app is protected by the next run — for a deployed app it is, for an undeployed one the card says so. Still open, deliberately: R-213, R-202, R-214, R-235.
233 lines
17 KiB
Markdown
233 lines
17 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
||
|
||
**Updated 2026-08-06.**
|
||
|
||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
|
||
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
|
||
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
|
||
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
|
||
> items rather than extend it.
|
||
|
||
## What works right now
|
||
|
||
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
|
||
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
|
||
home network, and open apps from a launcher or a shared link. Backups run on their own to three
|
||
places — the machine's drive, a second drive, and an encrypted off-site copy.
|
||
|
||
**The backup promise is proved. The recovery JOURNEY is not.** On 5 August we built a brand-new
|
||
machine from the published disc, gave it three marked files, destroyed it, and tried to get them back
|
||
**the way a household would** — no shortcuts, no command line. The files came back **byte for byte
|
||
identical**, all three, including one with Hungarian accents in its name. But **the journey needed us
|
||
four times**, and the very first thing the machine did was tell the customer their correct recovery
|
||
code was wrong. *(CAMPAIGN 11)*
|
||
|
||
## What's broken
|
||
|
||
- **A backup could report „✓ Rendben" while quietly leaving out an app the customer had just
|
||
chosen — FIXED 6 August.** Two things were wrong and only one had been guessed at. The machine's
|
||
own rule said *"a warning beside a success is read as a success"*, and it applied it to an app
|
||
missing a *folder* but not to an app left out *entirely*; that now counts too. And the real cause of
|
||
the case we saw: pressing „Távoli mentés most" while a backup was already running answered
|
||
„elindult" and then showed the **previous** run's green tick — so the customer read it as covering
|
||
their new app. It did not, and the restore refused minutes later. The button now says plainly that
|
||
it did not start anything.
|
||
|
||
- **A machine installed today still gets the older in-house service, so it cannot open a recovery
|
||
package until you approve the newer one.** It is no longer *lied to* — it says plainly that the
|
||
machine cannot do this yet — but **approving the new service is one click from you**, and until then
|
||
such a machine also gets the cautious "we do not know why" wording rather than the helpful one.
|
||
*(R-216, R-223, R-224)*
|
||
- **Three things a rebuilt machine still cannot do by itself.** Its owner cannot re-attach their own
|
||
drives, so no app can be put back on its data; it cannot create a new recovery code at all; and the
|
||
screen at the machine itself never stops showing a stale pairing code. Each is understood, measured
|
||
and written down — none is fixed yet. *(R-220, R-221, R-214)*
|
||
- **Rebuilding a machine still throws away its off-site backup HISTORY.** The machine invents the key
|
||
that encrypts its own off-site backups, and a rebuilt machine invents a new one. **The good news:
|
||
the old key really is kept now — we proved it on a real machine today, for the first time**, and a
|
||
changed key raises an alarm the same day. *(R-193, R-198)*
|
||
- **The kept older backups cannot be opened — by anyone.** We keep the previous sealed package and
|
||
there is no way to open it. The screens now say exactly that and stop. **One place still promises
|
||
otherwise**: the older-backups card says they "may be restorable later with the matching code",
|
||
which is not true today. *(R-222, R-202)*
|
||
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
|
||
stopgap. *(R-95, R-87)*
|
||
|
||
## Can a household get their data back on their own? Asked again today — still no, but nearer
|
||
|
||
We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it —
|
||
guest and both drives, as a hardware loss would — and tried to get them back the way a household
|
||
would. *(R-201, the re-walk)*
|
||
|
||
**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
|
||
Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
|
||
backup, in **23 seconds**, through the customer's own restore screen.
|
||
|
||
**And much of the journey now works.** The machine showed the recovery screen **without being asked**,
|
||
told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
|
||
code, and accepted the real code first time. The emailed claim code worked first try.
|
||
|
||
**But it still needed us twice**, and a household has neither hand:
|
||
|
||
- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
|
||
will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
|
||
can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
|
||
only half of it was)*
|
||
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
|
||
back, so the restore screen stays empty. *(R-220)*
|
||
|
||
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
|
||
|
||
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
|
||
built and published but not approved for new machines**. We installed them by hand for this test. So
|
||
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
|
||
|
||
## What we fixed this morning, and what it did not fix
|
||
|
||
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
|
||
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
|
||
cleared itself. What it found was that **the machine blamed the customer for failures that were not
|
||
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
|
||
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
|
||
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
|
||
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
|
||
|
||
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
|
||
remaining dead ends are different ones.
|
||
|
||
## What shipped recently
|
||
|
||
- **2026-08-06 (later)** — **The page that says which machine may be wrecked had one wrong sentence,
|
||
and it has been checked against the machine rather than corrected from memory.** It claimed the
|
||
spare HP box keeps no off-site copy, and used that as the reason to aim risky backup tests at the
|
||
other one. The HP box **does** keep an off-site copy — two of them, the newest from 4 August, in
|
||
its own space on the off-site server. The sentence was true when it was written and went out of
|
||
date on 23 July. It matters because a belief that a machine holds nothing is how a destructive test
|
||
lands on one that does. Also today: the notes the assistant keeps had three statements that were
|
||
simply untrue (including that the demo boxes were still away) — those three are fixed, the rest are
|
||
now flagged automatically rather than rewritten by hand. *(R-229, R-230)*
|
||
- **2026-08-06** — **The assistant's own working notes were on one machine with no copy anywhere.**
|
||
Everything the assistant has learned about this system over months — where things live, which traps
|
||
cost us an incident — sat in a folder on the build machine that no backup touched and no repository
|
||
held. It is now included in that machine's nightly backup. **Two things you should know before
|
||
treating that as solved:** the copy lands on the **same physical disk** as the original, so it
|
||
survives a mistake but not a dead drive, and the build machine's backups have **no off-site copy at
|
||
all**. **A full survey of that machine's backup, done the same day, confirmed both and found two
|
||
more things worth knowing.** The good news first: it has run every night without missing a set, and
|
||
we pulled a file back out of it and checked it matched the original exactly — the first time that
|
||
has ever been demonstrated. The rest: **if a backup ever fails, nobody is told** — the alert was
|
||
configured but never given anywhere to send to — and **every copy it makes stays inside that one
|
||
box**, so it survives any single disk dying but not the room. Nothing was changed; the survey was
|
||
read-only and the decisions are yours. Also tidied the same day: a quarter of those notes had become unreachable — filed but listed
|
||
nowhere, so nothing would ever read them — and the instructions the assistant reads at the start of
|
||
every session were cut roughly in half, with a check added so they cannot quietly grow back. Nothing
|
||
was deleted. *(R-229, R-230, R-231)*
|
||
- **2026-08-05 (late)** — **New machines now get current software again — and the disc image had been
|
||
quietly un-updatable for days.** Approving the newer in-house service turned out to be impossible on
|
||
its own: the system correctly refuses to publish a set where the pre-built machine image is older
|
||
than the software the fleet already runs, and that image had been behind since late July. **So the
|
||
image was rebuilt and both were published together.** A machine installed from now on lands on
|
||
current software and can open a recovery package on day one. *(R-223)*
|
||
|
||
- **2026-08-05 (evening)** — **Four sentences where there was one, and none of them blames you.**
|
||
"We did not accept your recovery code" used to appear when the code was wrong, when the machine
|
||
could not ask, when the store could not be read, and when the customer held the code for an older
|
||
backup we still keep. Only the first is the customer's doing. Also today: succeeding at recovery no
|
||
longer switches off the machine's own request for the thing it still needs; the screen now finishes
|
||
the job and shows what is in the backups instead of promising a list it could never produce; and the
|
||
recovery page can no longer be reached on a machine that never had backups.
|
||
*(R-216, R-217, R-218, R-219, R-222, R-215)*
|
||
|
||
- **2026-08-05** — **The recovery screen: a customer whose machine was rebuilt is now told, and shown
|
||
how.** Until today they had everything needed to get their data back and no way to find out — the
|
||
only route was a command line. The screen unlocks the backups and lists what is in them; it does
|
||
**not** restore anything, because unlocking and restoring are two different decisions and mixing
|
||
them would turn one clear moment into a wizard. Three ways out, none of them a dismiss button — and
|
||
the „most nem" option keeps the route to the data permanently visible in the backups area, because
|
||
a notice someone clicks past once is a notice that never happened. *(R-193)*
|
||
|
||
- **2026-08-05** — **The orphaned backups are deleted — and the list you were given was wrong, which
|
||
is why you were asked again.** You had approved "about 1.2 GB in two set-aside stores". Measured
|
||
before touching anything: there were **three** set-aside stores totalling **~1.45 GB** — and the
|
||
thing that was exactly 1.2 GB was demo-felhom's **live** store. Matching on the size would have
|
||
deleted a working backup. With the corrected list confirmed, all three were removed and both live
|
||
stores left alone; a real off-site backup ran successfully straight afterwards to prove nothing
|
||
working had been caught. *(R-212)*
|
||
|
||
- **2026-08-05** — **A rebuilt machine now asks for its storage credential, and the hub gives it back.**
|
||
The machine says plainly what it needs — it can tell it has been rebuilt, because its data area is
|
||
empty *and* the hub is holding a sealed recovery package for it — instead of leaving the hub to guess
|
||
from a silence that has four possible meanings. The hub waits long enough to be sure it is not a
|
||
restart, then **re-uses the credential it already holds** before creating a new one at the storage
|
||
provider. **The recovery ceremony stays manual, deliberately:** a credential can be replaced, your
|
||
recovery code cannot. *(R-204 item 4, R-193)*
|
||
|
||
- **2026-08-05** — **The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer
|
||
is unstuck.** The cause was leftover working data from building our own software — 157 GB of it,
|
||
growing about 5 GB a day, which nothing was allowed to delete. **148 GB came back in 86 seconds.**
|
||
The storage layer had already stopped accepting new copies of any volume onto that disk; that is
|
||
fixed the same day. A **30 GB ceiling** is now in place and was **proved to work by deliberately
|
||
overfilling it and watching it evict** — not by assuming the setting took. Two things that failed
|
||
quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved),
|
||
and the weekly cleanup is still forbidden from touching the thing that grows (next session).
|
||
*(R-205 … R-211)*
|
||
- **2026-08-05** — **The build data now lives on the second SSD, moved with nothing lost.** All 345
|
||
images, every saved volume and both development databases came through identical — checked before
|
||
the original was touched and again afterwards, and confirmed by running a real build on the moved
|
||
copy. **The server's own services never went down:** Gitea, the registry, the hub and the backup
|
||
system run on a separate system and stayed up throughout. The second SSD now also **reserves 80 GB**
|
||
for this, so the storage layer can no longer quietly claim the space and repeat what happened to the
|
||
first disk. The old copy is kept as the way back until the machine next restarts. *(R-209)*
|
||
- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time;
|
||
a credential re-issue no longer blocks off-site backups on a healthy machine; and the default
|
||
restore no longer quietly returns the wrong thing. *(R-204 items 1–3, R-196)*
|
||
- **2026-08-04 (night)** — **The drill PASSED**, end to end, on real hardware. *(R-201)*
|
||
- **2026-08-04** — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup
|
||
look in the same directory, and a backup that misses a folder marked essential reports *incomplete*
|
||
instead of success. *(R-203)*
|
||
- **2026-08-04** — The hub now keeps the off-site backup key when a machine re-seals, instead of only
|
||
the whole-machine one, and a machine can fetch its own sealed package back. *(R-198, R-199)*
|
||
- **2026-08-04** — The daily false alarm about David is gone; a vanished permission now repairs itself
|
||
**and says it had to**; the weekly off-site backup stopped reporting failure after a successful
|
||
upload. *(R-195, R-190, R-191)*
|
||
|
||
## What we're working on
|
||
|
||
- **Next: the retention proof.** The hub keeping the old sealed key when a machine re-seals is the one
|
||
remaining link that has **never run outside a test**. Proving it needs a second deliberate wipe on
|
||
the spare demo machine, and it is its own procedure. *(R-198)*
|
||
- **Then:** the orphaned-backup deletion you asked for (below), and the off-site copy the machine can
|
||
still erase. *(R-193, R-95)*
|
||
|
||
## Waiting on you
|
||
|
||
|
||
- **One thing to read after the machine next restarts — and nothing to do until then.** You told me
|
||
not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been
|
||
through a restart. It works right now and nothing was lost, but a restart is the one test that
|
||
matters for this kind of change, and it has not happened. **I made it check itself:** whenever the
|
||
machine next starts, for any reason, it writes a plain PASS or FAIL line to
|
||
`/var/log/felhom-store-postboot-check.log`. **If it says PASS, the old copy can be deleted and 34 GB
|
||
comes back.** Until then I have deliberately kept that old copy, which is the only quick way back —
|
||
it is why the disk sits at 54% rather than lower. *(R-209a)*
|
||
- **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and
|
||
62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old
|
||
(controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing
|
||
was deleted.** Worth knowing before you spend time on it: they only account for about 27 GB against
|
||
199 GB now free, so this is about clutter, not space. *(R-210)*
|
||
- **Nothing.**
|
||
- **The recovery screen you described has been priced, and it can be built.** A freshly installed
|
||
machine that finds a sealed package waiting should say so, offer a box for the recovery code, and
|
||
show what would come back before doing anything. One thing to weigh, deliberately not decided: that
|
||
screen is reachable by anyone with the household's dashboard password, and the preview reveals
|
||
backup dates and app names. *(R-193)*
|
||
- **The orphaned backups on the storage box — you said delete, and it is still owed.** About 1.2 GB
|
||
across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its
|
||
own session rather than riding along with other work. *(R-193)*
|
||
- *(decided 4 Aug)* You chose **not** to keep a copy of the backup key on the Proxmox host, which
|
||
makes the customer's own recovery code the only route back from a rebuild. *(R-193)*
|
||
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
|
||
session log; nothing suggests anyone else saw it. *(R-132)*
|
||
- **One small question, not urgent.** The automatic version check cannot see which version you have
|
||
told machines to install, only which ones exist. *(R-184)*
|