0c4411e54b
gates / gates (push) Successful in 9s
Asked Campaign 11 Phase 1's question a second time, on the fixed build, on a
NEW appliance (VM 322, customer rewalk). The Campaign 11 venue was untouched.
THE DATA: PASS. All three sentinels byte-identical out of the pre-destruction
snapshot a7bc23bd in 23s through the customer's own restore flow — including a
12 MB binary and an accented Hungarian filename whose NAME BYTES are identical
too (verified as hex, not as rendered text).
THE JOURNEY: FAIL, two dead ends against Phase 1's four.
1. R-218's CONSUME half. The hub re-staged the credential at 11:44:57 saying
'the box re-consumes on its next cycle'; a full cycle ran at 11:55:46/54
(with a positive control that it ran) and it did not. A census of the
customer-reachable actions found none that fetches it. Only a command line
INSIDE THE GUEST moved it — 18s, confirming nothing was wrong with the
credential, target or key: only the trigger. R-218's row said SHIPPED and
over-claimed; it is corrected to REOPENED for the consume half.
2. R-220. Drives still unenrollable after a rebuild, needing a Proxmox-host
unmount; without it no app redeploys and the restore page stays empty.
Unaided RTO STILL UNDEFINED. Attended: +45s key placed, +24m12s tier up,
+30m13s data verified. The 30m must not be quoted as the customer number.
What passed and is new: the recovery screen appeared WITHOUT being sought,
answered all three questions with a seal date matching the hub exactly, the
emailed reset code worked first try, the unlock was a real 1.528s unseal, and
R-225's fix was seen working in the wild (unknown, not a false zero).
R-216 part 4 reproduced live: the reinstall downgraded the hand-installed agent
0.126.0 -> 0.125.0.
DELIVERY GAP recorded as owed and NOT conflated with the journey: a fresh
install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions,
neither carrying the fixes — installed by hand. Nothing was vouched.
Capability map row STAYS FAIL. Campaign 11 doc gets a dated ADDENDUM, not a
rewrite.
224 lines
17 KiB
Markdown
224 lines
17 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
||
|
||
**Updated 2026-08-06.**
|
||
|
||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
|
||
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
|
||
> which is technical state written for Claude Code — keep the two separate. **Maintenance:** update
|
||
> at the end of every session in which something shipped, broke, or was decided. One screen; cut
|
||
> items rather than extend it.
|
||
|
||
## What works right now
|
||
|
||
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
|
||
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
|
||
home network, and open apps from a launcher or a shared link. Backups run on their own to three
|
||
places — the machine's drive, a second drive, and an encrypted off-site copy.
|
||
|
||
**The backup promise is proved. The recovery JOURNEY is not.** On 5 August we built a brand-new
|
||
machine from the published disc, gave it three marked files, destroyed it, and tried to get them back
|
||
**the way a household would** — no shortcuts, no command line. The files came back **byte for byte
|
||
identical**, all three, including one with Hungarian accents in its name. But **the journey needed us
|
||
four times**, and the very first thing the machine did was tell the customer their correct recovery
|
||
code was wrong. *(CAMPAIGN 11)*
|
||
|
||
## What's broken
|
||
|
||
- **A machine installed today still gets the older in-house service, so it cannot open a recovery
|
||
package until you approve the newer one.** It is no longer *lied to* — it says plainly that the
|
||
machine cannot do this yet — but **approving the new service is one click from you**, and until then
|
||
such a machine also gets the cautious "we do not know why" wording rather than the helpful one.
|
||
*(R-216, R-223, R-224)*
|
||
- **Three things a rebuilt machine still cannot do by itself.** Its owner cannot re-attach their own
|
||
drives, so no app can be put back on its data; it cannot create a new recovery code at all; and the
|
||
screen at the machine itself never stops showing a stale pairing code. Each is understood, measured
|
||
and written down — none is fixed yet. *(R-220, R-221, R-214)*
|
||
- **Rebuilding a machine still throws away its off-site backup HISTORY.** The machine invents the key
|
||
that encrypts its own off-site backups, and a rebuilt machine invents a new one. **The good news:
|
||
the old key really is kept now — we proved it on a real machine today, for the first time**, and a
|
||
changed key raises an alarm the same day. *(R-193, R-198)*
|
||
- **The kept older backups cannot be opened — by anyone.** We keep the previous sealed package and
|
||
there is no way to open it. The screens now say exactly that and stop. **One place still promises
|
||
otherwise**: the older-backups card says they "may be restorable later with the matching code",
|
||
which is not true today. *(R-222, R-202)*
|
||
- **The off-site copy can be erased by the machine that made it.** A daily snapshot is armed as a
|
||
stopgap. *(R-95, R-87)*
|
||
|
||
## Can a household get their data back on their own? Asked again today — still no, but nearer
|
||
|
||
We built a **brand-new machine** from the published disc, gave it three marked files, destroyed it —
|
||
guest and both drives, as a hardware loss would — and tried to get them back the way a household
|
||
would. *(R-201, the re-walk)*
|
||
|
||
**The files came back perfectly.** All three, byte for byte, including a 12 MB file and one whose
|
||
Hungarian accented filename came back **letter-for-letter identical**. Out of the pre-destruction
|
||
backup, in **23 seconds**, through the customer's own restore screen.
|
||
|
||
**And much of the journey now works.** The machine showed the recovery screen **without being asked**,
|
||
told the customer what was waiting and when it was sealed, said plainly that nobody can replace a lost
|
||
code, and accepted the real code first time. The emailed claim code worked first try.
|
||
|
||
**But it still needed us twice**, and a household has neither hand:
|
||
|
||
- **The machine never picks up its own storage connection.** Our hub hands it over and says "the box
|
||
will collect this on its next cycle" — the cycle came and went and it did not. Nothing the customer
|
||
can click fixes it; it took a command inside the machine. *(R-218 — we had recorded this as fixed;
|
||
only half of it was)*
|
||
- **A rebuilt machine still cannot re-attach its own drives** — and without them no app can be put
|
||
back, so the restore screen stays empty. *(R-220)*
|
||
|
||
**Two dead ends, down from four.** The verdict stays **FAILED** until a walk needs us zero times.
|
||
|
||
**One thing to decide.** A machine installed today still gets the older software — **the fixes are
|
||
built and published but not approved for new machines**. We installed them by hand for this test. So
|
||
this proves the journey works on the fixed build; it does **not** prove a customer would receive it.
|
||
|
||
## What we fixed this morning, and what it did not fix
|
||
|
||
Overnight we tried to break the recovery journey with eleven faults and then left the machine alone
|
||
for a full cycle. Nothing lost a byte; the set-aside backups really were kept; the alarm fired and
|
||
cleared itself. What it found was that **the machine blamed the customer for failures that were not
|
||
theirs** — our hub being unreachable came back as "your recovery code is wrong", in three hundredths
|
||
of a second, without the machine even trying. **That is fixed and deployed** *(R-224, R-226)*, along
|
||
with three smaller truths: "0 snapshots" where the answer is "we have not looked yet" *(R-225)*, a raw
|
||
English error mid-recovery *(R-227)*, and set-aside backups that had become invisible *(R-228)*.
|
||
|
||
**None of that shortened the journey**, which is why today's re-walk above still says FAILED — the two
|
||
remaining dead ends are different ones.
|
||
|
||
## What shipped recently
|
||
|
||
- **2026-08-06 (later)** — **The page that says which machine may be wrecked had one wrong sentence,
|
||
and it has been checked against the machine rather than corrected from memory.** It claimed the
|
||
spare HP box keeps no off-site copy, and used that as the reason to aim risky backup tests at the
|
||
other one. The HP box **does** keep an off-site copy — two of them, the newest from 4 August, in
|
||
its own space on the off-site server. The sentence was true when it was written and went out of
|
||
date on 23 July. It matters because a belief that a machine holds nothing is how a destructive test
|
||
lands on one that does. Also today: the notes the assistant keeps had three statements that were
|
||
simply untrue (including that the demo boxes were still away) — those three are fixed, the rest are
|
||
now flagged automatically rather than rewritten by hand. *(R-229, R-230)*
|
||
- **2026-08-06** — **The assistant's own working notes were on one machine with no copy anywhere.**
|
||
Everything the assistant has learned about this system over months — where things live, which traps
|
||
cost us an incident — sat in a folder on the build machine that no backup touched and no repository
|
||
held. It is now included in that machine's nightly backup. **Two things you should know before
|
||
treating that as solved:** the copy lands on the **same physical disk** as the original, so it
|
||
survives a mistake but not a dead drive, and the build machine's backups have **no off-site copy at
|
||
all**. **A full survey of that machine's backup, done the same day, confirmed both and found two
|
||
more things worth knowing.** The good news first: it has run every night without missing a set, and
|
||
we pulled a file back out of it and checked it matched the original exactly — the first time that
|
||
has ever been demonstrated. The rest: **if a backup ever fails, nobody is told** — the alert was
|
||
configured but never given anywhere to send to — and **every copy it makes stays inside that one
|
||
box**, so it survives any single disk dying but not the room. Nothing was changed; the survey was
|
||
read-only and the decisions are yours. Also tidied the same day: a quarter of those notes had become unreachable — filed but listed
|
||
nowhere, so nothing would ever read them — and the instructions the assistant reads at the start of
|
||
every session were cut roughly in half, with a check added so they cannot quietly grow back. Nothing
|
||
was deleted. *(R-229, R-230, R-231)*
|
||
- **2026-08-05 (late)** — **New machines now get current software again — and the disc image had been
|
||
quietly un-updatable for days.** Approving the newer in-house service turned out to be impossible on
|
||
its own: the system correctly refuses to publish a set where the pre-built machine image is older
|
||
than the software the fleet already runs, and that image had been behind since late July. **So the
|
||
image was rebuilt and both were published together.** A machine installed from now on lands on
|
||
current software and can open a recovery package on day one. *(R-223)*
|
||
|
||
- **2026-08-05 (evening)** — **Four sentences where there was one, and none of them blames you.**
|
||
"We did not accept your recovery code" used to appear when the code was wrong, when the machine
|
||
could not ask, when the store could not be read, and when the customer held the code for an older
|
||
backup we still keep. Only the first is the customer's doing. Also today: succeeding at recovery no
|
||
longer switches off the machine's own request for the thing it still needs; the screen now finishes
|
||
the job and shows what is in the backups instead of promising a list it could never produce; and the
|
||
recovery page can no longer be reached on a machine that never had backups.
|
||
*(R-216, R-217, R-218, R-219, R-222, R-215)*
|
||
|
||
- **2026-08-05** — **The recovery screen: a customer whose machine was rebuilt is now told, and shown
|
||
how.** Until today they had everything needed to get their data back and no way to find out — the
|
||
only route was a command line. The screen unlocks the backups and lists what is in them; it does
|
||
**not** restore anything, because unlocking and restoring are two different decisions and mixing
|
||
them would turn one clear moment into a wizard. Three ways out, none of them a dismiss button — and
|
||
the „most nem" option keeps the route to the data permanently visible in the backups area, because
|
||
a notice someone clicks past once is a notice that never happened. *(R-193)*
|
||
|
||
- **2026-08-05** — **The orphaned backups are deleted — and the list you were given was wrong, which
|
||
is why you were asked again.** You had approved "about 1.2 GB in two set-aside stores". Measured
|
||
before touching anything: there were **three** set-aside stores totalling **~1.45 GB** — and the
|
||
thing that was exactly 1.2 GB was demo-felhom's **live** store. Matching on the size would have
|
||
deleted a working backup. With the corrected list confirmed, all three were removed and both live
|
||
stores left alone; a real off-site backup ran successfully straight afterwards to prove nothing
|
||
working had been caught. *(R-212)*
|
||
|
||
- **2026-08-05** — **A rebuilt machine now asks for its storage credential, and the hub gives it back.**
|
||
The machine says plainly what it needs — it can tell it has been rebuilt, because its data area is
|
||
empty *and* the hub is holding a sealed recovery package for it — instead of leaving the hub to guess
|
||
from a silence that has four possible meanings. The hub waits long enough to be sure it is not a
|
||
restart, then **re-uses the credential it already holds** before creating a new one at the storage
|
||
provider. **The recovery ceremony stays manual, deliberately:** a credential can be replaced, your
|
||
recovery code cannot. *(R-204 item 4, R-193)*
|
||
|
||
- **2026-08-05** — **The DooPlex server's disk is out of danger: 86% full → 54%, and the storage layer
|
||
is unstuck.** The cause was leftover working data from building our own software — 157 GB of it,
|
||
growing about 5 GB a day, which nothing was allowed to delete. **148 GB came back in 86 seconds.**
|
||
The storage layer had already stopped accepting new copies of any volume onto that disk; that is
|
||
fixed the same day. A **30 GB ceiling** is now in place and was **proved to work by deliberately
|
||
overfilling it and watching it evict** — not by assuming the setting took. Two things that failed
|
||
quietly around it: the warning meant to catch exactly this **could never fire** (fixed and proved),
|
||
and the weekly cleanup is still forbidden from touching the thing that grows (next session).
|
||
*(R-205 … R-211)*
|
||
- **2026-08-05** — **The build data now lives on the second SSD, moved with nothing lost.** All 345
|
||
images, every saved volume and both development databases came through identical — checked before
|
||
the original was touched and again afterwards, and confirmed by running a real build on the moved
|
||
copy. **The server's own services never went down:** Gitea, the registry, the hub and the backup
|
||
system run on a separate system and stayed up throughout. The second SSD now also **reserves 80 GB**
|
||
for this, so the storage layer can no longer quietly claim the space and repeat what happened to the
|
||
first disk. The old copy is kept as the way back until the machine next restarts. *(R-209)*
|
||
- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time;
|
||
a credential re-issue no longer blocks off-site backups on a healthy machine; and the default
|
||
restore no longer quietly returns the wrong thing. *(R-204 items 1–3, R-196)*
|
||
- **2026-08-04 (night)** — **The drill PASSED**, end to end, on real hardware. *(R-201)*
|
||
- **2026-08-04** — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup
|
||
look in the same directory, and a backup that misses a folder marked essential reports *incomplete*
|
||
instead of success. *(R-203)*
|
||
- **2026-08-04** — The hub now keeps the off-site backup key when a machine re-seals, instead of only
|
||
the whole-machine one, and a machine can fetch its own sealed package back. *(R-198, R-199)*
|
||
- **2026-08-04** — The daily false alarm about David is gone; a vanished permission now repairs itself
|
||
**and says it had to**; the weekly off-site backup stopped reporting failure after a successful
|
||
upload. *(R-195, R-190, R-191)*
|
||
|
||
## What we're working on
|
||
|
||
- **Next: the retention proof.** The hub keeping the old sealed key when a machine re-seals is the one
|
||
remaining link that has **never run outside a test**. Proving it needs a second deliberate wipe on
|
||
the spare demo machine, and it is its own procedure. *(R-198)*
|
||
- **Then:** the orphaned-backup deletion you asked for (below), and the off-site copy the machine can
|
||
still erase. *(R-193, R-95)*
|
||
|
||
## Waiting on you
|
||
|
||
|
||
- **One thing to read after the machine next restarts — and nothing to do until then.** You told me
|
||
not to restart DooPlex, so I did not, and the move to the second SSD has therefore never been
|
||
through a restart. It works right now and nothing was lost, but a restart is the one test that
|
||
matters for this kind of change, and it has not happened. **I made it check itself:** whenever the
|
||
machine next starts, for any reason, it writes a plain PASS or FAIL line to
|
||
`/var/log/felhom-store-postboot-check.log`. **If it says PASS, the old copy can be deleted and 34 GB
|
||
comes back.** Until then I have deliberately kept that old copy, which is the only quick way back —
|
||
it is why the disk sits at 54% rather than lower. *(R-209a)*
|
||
- **One list to rule on: 193 old images that exist only on this machine.** 131 controller versions and
|
||
62 hub versions are not in the registry, so they cannot be re-downloaded — all of them old
|
||
(controller up to 0.135.0, hub up to 0.57.0; everything newer is safely in the registry). **Nothing
|
||
was deleted.** Worth knowing before you spend time on it: they only account for about 27 GB against
|
||
199 GB now free, so this is about clutter, not space. *(R-210)*
|
||
- **Nothing.**
|
||
- **The recovery screen you described has been priced, and it can be built.** A freshly installed
|
||
machine that finds a sealed package waiting should say so, offer a box for the recovery code, and
|
||
show what would come back before doing anything. One thing to weigh, deliberately not decided: that
|
||
screen is reachable by anyone with the household's dashboard password, and the preview reveals
|
||
backup dates and app names. *(R-193)*
|
||
- **The orphaned backups on the storage box — you said delete, and it is still owed.** About 1.2 GB
|
||
across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its
|
||
own session rather than riding along with other work. *(R-193)*
|
||
- *(decided 4 Aug)* You chose **not** to keep a copy of the backup key on the Proxmox host, which
|
||
makes the customer's own recovery code the only route back from a rebuild. *(R-193)*
|
||
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
|
||
session log; nothing suggests anyone else saw it. *(R-132)*
|
||
- **One small question, not urgent.** The automatic version check cannot see which version you have
|
||
told machines to install, only which ones exist. *(R-184)*
|