docs: R-196 closed, R-204 items 1-3 closed, item 4 open (R-193)
gates / gates (push) Successful in 7s
gates / gates (push) Successful in 7s
- OPEN-ITEMS: R-196 CLOSED; R-204 items 1-3 CLOSED with item 4 named and its dependency stated. Header restates that R-202, the 1.2 GB ciphertext deletion and R-198's still-unit-proven retention all REMAIN OPEN. - capability map: the recovery row keeps its 'with a person present' qualifier, names which crutch remains, and cites the three now gone. - 07-backup-architecture: new 7.0 - what a customer can and cannot do ALONE, the four steps in a table with status. This is the section a future reader will use to answer that question. - CONTEXT: standing ruling S-32, superseding S-31 steps 2-5. - STATUS: rewritten to one screen per its own header; removes a corrupted half-overwritten section left from the drill session. - ROADMAP: R-196 and R-204 collapsed.
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-08-04.**
|
||||
**Updated 2026-08-05.**
|
||||
|
||||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this
|
||||
> page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**,
|
||||
@@ -13,179 +13,78 @@
|
||||
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer,
|
||||
who sets their own password. They install apps from a catalogue of fifty-three, share files over the
|
||||
home network, and open apps from a launcher or a shared link. Backups run on their own to three
|
||||
places — the machine's drive, a second drive, and an encrypted off-site copy — and a customer can
|
||||
restore files and app data from the drive alone. Apps come back after a power cut: hard-reset the demo
|
||||
box six times, everything returned every time, and an app switched off deliberately stayed off.
|
||||
Proven end to end on real hardware.
|
||||
places — the machine's drive, a second drive, and an encrypted off-site copy.
|
||||
|
||||
**And the whole backup promise is now proved.** On 4 August we destroyed a machine on purpose and
|
||||
deleted a marked file from its disk. Using the recovery code you saved: the backup key came back
|
||||
**identical, character for character**; the existing off-site store **opened** rather than starting
|
||||
over; and the file was restored **byte for byte identical**. *(R-201)*
|
||||
|
||||
## What's broken
|
||||
|
||||
- **A customer still cannot do that recovery alone — but only one step is left.** Getting from "the
|
||||
key is recoverable" to "the file is back" took four steps that appeared in no instructions. **Three
|
||||
are fixed today** *(R-204)*: the local reset-code tool works on the first try instead of needing the
|
||||
controller restarted; re-issuing the storage credential no longer falsely marks the recovery key
|
||||
"stale" (which used to stop every off-site backup and invite the one act that would have destroyed
|
||||
the recovered key); and the everyday restore now **says in plain Hungarian that it returned the
|
||||
app's settings and database and not your documents**, and names the button that does. **The step
|
||||
that remains is the first one:** a rebuilt machine cannot get a storage credential by itself,
|
||||
because the one-time password was used up by its predecessor — so you still have to press
|
||||
Re-issue. That is a design decision waiting on you, below. *(R-193)*
|
||||
- **Rebuilding a machine still throws away its off-site backup HISTORY.** The machine invents the key
|
||||
that encrypts its own off-site backups, and a rebuilt machine invents a brand-new one, so the old
|
||||
backups can no longer be opened by the machine that wrote them. Both demo machines did this on 3–4
|
||||
August — demo-hp lost access to 15 backups (41 MB), demo-felhom to 36 (1.14 GB), and **the second one
|
||||
happened in silence for thirteen hours**. From today that silence is fixed (below) and the old key
|
||||
survives the ceremony (below), but **the rebuild itself still starts a fresh history**, and getting
|
||||
the old one back is a path nobody has walked. *(R-193)*
|
||||
- *(fixed 4 Aug)* ~~The safety net we believed was under all of this is not there.~~ **The central hub
|
||||
was keeping the wrong key, and the fix shipped the same day.** When a machine re-seals its backups
|
||||
under a new recovery code, the hub keeps the old sealed package — but it was keeping the one for the
|
||||
whole-machine backups and **not** the one for the off-site file backups, which is the whole point. And
|
||||
because the swap happens in one motion, **the act that destroyed the old key was the recovery-code
|
||||
ceremony itself** — the exact thing a rebuilt machine asks its customer to do, on a screen promising
|
||||
the old backups stay recoverable. Both keys are now kept. **What this cannot undo:** anything
|
||||
superseded before today is gone for good, which includes both demo machines' pre-4-August keys — so
|
||||
those 51 orphaned backups were beyond reach even if the recovery codes had been kept. *(R-198)*
|
||||
- **IT WORKS. A customer's file survived a machine being destroyed and came back — proved last night on
|
||||
real hardware, end to end.** We wiped the HP machine's controller data on purpose and deleted the
|
||||
marked file from its disk. Using the recovery code you saved: the backup key came back **identical,
|
||||
character for character**; the existing off-site store **opened** rather than starting over (the same
|
||||
three backups, the same 42 026 bytes as before the wipe); and the file was restored **byte for byte
|
||||
identical**. That is the whole backup promise, demonstrated for the first time. *(R-201)*
|
||||
- **But a customer could not have done it alone, and that is now the open work.** Getting from "the key
|
||||
is recoverable" to "the file is back" took four steps that appear in no instructions: re-issuing the
|
||||
storage credential; re-claiming the machine (whose local reset-code tool **does not work until the
|
||||
controller is restarted** — two attempts failed before we spotted it); manually clearing a "stale"
|
||||
mark that the re-issue sets even though the key never changed; and choosing the *full* restore,
|
||||
because the default one returns the app's settings and **not** the customer's documents — with
|
||||
nothing saying so. **The capability is real; the experience is not built yet.** *(R-204)*
|
||||
that encrypts its own off-site backups, and a rebuilt machine invents a brand-new one. Both demo
|
||||
machines did this on 3–4 August — 51 backups (~1.2 GB) between them. The old key now survives the
|
||||
recovery ceremony, and a changed key now raises an alarm the same day, but **the rebuild itself
|
||||
still starts a fresh history**. *(R-193)*
|
||||
- **One screen still tells the customer something we cannot yet promise.** The „elárvult tároló" card
|
||||
says the old backups may later be restorable with the matching recovery code. From today that is true
|
||||
for machines that re-seal from now on and **false for anything already orphaned** — and the machine
|
||||
cannot tell which case it is in without the recovery path above. We deliberately did **not** patch the
|
||||
sentence: a conditional promise that can still be wrong is worse on that screen than a vague one.
|
||||
*(R-202)*
|
||||
- *(fixed 4 Aug)* ~~The weekly off-site backup reports FAILED although it worked.~~ It uploaded fine
|
||||
and then tripped on a tidy-up step it is deliberately not allowed to perform. The machine no longer
|
||||
asks — tidying up is the endpoint's job, and **that was checked first**: the endpoint has been doing
|
||||
it every day since 27 July. Both machines and all future ones. *(R-191)*
|
||||
says the old backups may later be restorable with the matching recovery code. That is true for
|
||||
machines that re-seal from now on and **false for anything already orphaned** — and the machine
|
||||
cannot tell which case it is in. We deliberately did **not** patch the sentence: a conditional
|
||||
promise that can still be wrong is worse there than a vague one. *(R-202)*
|
||||
- **The off-site copy can be erased by the machine that made it.** The credential that writes it can
|
||||
also delete it. A daily snapshot is armed as a stopgap.
|
||||
*(R-95, R-87)*
|
||||
also delete it. A daily snapshot is armed as a stopgap. *(R-95, R-87)*
|
||||
|
||||
## What shipped recently
|
||||
|
||||
- **The daily error email about David is gone.** David's record was created when he said he would try
|
||||
Felhom; no machine was ever attached to it, so nothing has ever been expected of it — yet every
|
||||
morning at five the system reported that a database backup had not happened. **A customer with no
|
||||
machine ever attached now produces no alarm.** A machine that *was* attached and went quiet still
|
||||
does, because that is a real fault, and there is a test whose whole job is to make sure a future
|
||||
change cannot take that away. David's record was not touched — the record was right; the alarm was
|
||||
wrong. **The first quiet morning will be 5 August, so this is not confirmed yet.** *(R-195)*
|
||||
|
||||
- **A permission the machine needs can now repair itself — and it tells you it had to.** Yesterday we
|
||||
found a permission that vanished with no reinstall and no recorded action; the machine now notices,
|
||||
restores it through the same controlled helper that grants it in the first place, and **raises an
|
||||
alert saying it was missing and was restored**. The alert is the point: a permission that silently
|
||||
fixes itself is one that can be lost over and over and never investigated. Proven twice on a demo
|
||||
machine by taking the permission away by hand. *(R-190 — the repair; why it vanished is still open)*
|
||||
|
||||
- **The on-machine backup copy has now been proved to restore — by the machines themselves**, both of
|
||||
them, overnight and unasked (84 and 109 seconds). Every restore proof before this was of the
|
||||
off-site copy. *(R-86, R-185)*
|
||||
- **Three ways the alarm system was misreporting its own work — all fixed.** None of them ever risked
|
||||
data. **(1)** When the machine proved a backup restores, that result could vanish if the agent was
|
||||
restarted in the following quarter-hour — and yesterday's change made the gap a week rather than a
|
||||
day, because the machine correctly refuses to re-prove an archive it has already proven. It is now
|
||||
written to disk with the result and survives. This was caught happening, not predicted: a real
|
||||
14.5 GB off-site restore passed and left no record at all. **(2)** Every release had about a
|
||||
fifty-fifty chance of emailing you a failure for a release that worked; the version tag is now
|
||||
published after the binary, and a new check catches the opposite mistake so nothing is traded away.
|
||||
**(3)** A released binary can now be rebuilt by anyone and checked against the fingerprint you
|
||||
approve — until today, rebuilding produced different bytes. *(R-189, R-188, R-186)*
|
||||
|
||||
- **2026-08-05** — **Three of the four recovery crutches removed.** The reset code works first time;
|
||||
a credential re-issue no longer blocks off-site backups on a healthy machine; and the default
|
||||
restore no longer quietly returns the wrong thing. *(R-204 items 1–3, R-196)*
|
||||
- **2026-08-04 (night)** — **The drill PASSED**, end to end, on real hardware. *(R-201)*
|
||||
- **2026-08-04** — The folder-left-out-of-the-backup problem fixed both halves: the app and its backup
|
||||
look in the same directory, and a backup that misses a folder marked essential reports *incomplete*
|
||||
instead of success. *(R-203)*
|
||||
- **2026-08-04** — The hub now keeps the off-site backup key when a machine re-seals, instead of only
|
||||
the whole-machine one, and a machine can fetch its own sealed package back. *(R-198, R-199)*
|
||||
- **2026-08-04** — The daily false alarm about David is gone; a vanished permission now repairs itself
|
||||
**and says it had to**; the weekly off-site backup stopped reporting failure after a successful
|
||||
upload. *(R-195, R-190, R-191)*
|
||||
|
||||
## What we're working on
|
||||
|
||||
- **Now:** making that journey something a customer can actually follow — the four steps above, worst
|
||||
first: the restore default that silently returns the wrong thing. The HP machine is sitting
|
||||
mid-drill with the recovered key already on it; it needs re-claiming and one setting confirmed, then
|
||||
the backup runs and the file is restored. **Do not let it make a new recovery code** — that would
|
||||
The marked file now lands in
|
||||
the off-site backup, so there is finally something to recover. Everything else is already in place on
|
||||
the HP machine — the recovery code you saved, a working off-site store, the app and the file. Nothing
|
||||
has been wiped; that step waits for your go-ahead at a marked stop point. Both honesty fixes shipped today: a changed backup key now
|
||||
raises an alarm on the day, and the email that stated the opposite of what it measured now describes
|
||||
what it actually saw.
|
||||
- **Also tomorrow:** confirming what the 04:15 off-site run does on both machines. We expect it to
|
||||
refuse and ask, rather than start fresh — but that is a prediction from the code, not something
|
||||
anyone has watched happen.
|
||||
- **Next:** proving the off-site *app-data* copy can actually be restored — the one tier nothing
|
||||
tests unattended. Most of the machinery it needed arrived with the restore-test change below.
|
||||
*(R-87)*
|
||||
- **After:** the off-site copy that the machine making it can still erase. *(R-95)*
|
||||
- **Next: the retention proof.** The hub keeping the old sealed key when a machine re-seals is the one
|
||||
remaining link that has **never run outside a test**. Proving it needs a second deliberate wipe on
|
||||
the spare demo machine, and it is its own procedure. *(R-198)*
|
||||
- **Then:** the orphaned-backup deletion you asked for (below), and the off-site copy the machine can
|
||||
still erase. *(R-193, R-95)*
|
||||
|
||||
## Waiting on you
|
||||
|
||||
- *(decided 4 Aug — recorded here until the work it unblocks is done)* You chose **not** to keep a
|
||||
copy of the backup key on the Proxmox host. That is a coherent choice and it has a consequence worth
|
||||
restating: **the customer's own recovery code is now the only route back from a rebuild**, so the
|
||||
sealed package the hub keeps is the single copy. Today's fix is what makes that true rather than
|
||||
wishful — until this morning the ceremony destroyed it. *(R-193)*
|
||||
- **Your ruling about the recovery screen has been priced, and it can be built.** You said: a freshly
|
||||
installed machine that finds a sealed package waiting should say so loudly, offer the customer a box
|
||||
to type their recovery code into, and show what would come back before doing anything. **All of it is
|
||||
buildable, and one part is already free** — the hub is *already* telling every machine, on every
|
||||
check-in, that a sealed package exists and which key it covers, and the machine currently throws that
|
||||
message away. Showing a preview is also cheap: listing what is in an off-site store reads it without
|
||||
writing to it, so the customer can see how many backups, from when, and for which apps before
|
||||
committing. The real work is one new connection: the unsealing has to happen in the part that runs on
|
||||
the Proxmox host, because the customer-facing part cannot do it — the same crossing your recovery-code
|
||||
ceremony already makes in the other direction. **One thing for you to weigh, which we deliberately did
|
||||
not decide:** a screen that takes a recovery code and then shows what is in a backup store is reachable
|
||||
by anyone with the household's dashboard password, and the preview reveals backup dates and app names.
|
||||
*(R-193)*
|
||||
- **Should we run the proof?** Nothing about recovery has ever actually been done, so we designed the
|
||||
exercise: on the spare demo machine, make a recovery code and **keep it**, put a marked file into an
|
||||
app, back it up off-site, wipe the machine, reinstall, recover with the saved code, and check the
|
||||
marked file comes back byte for byte. About half a day, of which under an hour needs you. The failure
|
||||
we are watching for is precise: **if the recovered machine reports one backup instead of the ones we
|
||||
put there, it started a fresh history and the proof failed** — "the store opened" is not good enough.
|
||||
It costs that machine's contents if it goes wrong, which is acceptable there. **You said yes, after
|
||||
R-198 — and R-198 shipped this morning, so this is ready to run on your word.** *(R-201)*
|
||||
- **The one-shot credential decision — this is now the last thing between a customer and an unaided
|
||||
recovery.** A rebuilt machine has no storage credential of its own, so an operator must press
|
||||
Re-issue. Everything after that point is now self-service. Deciding how a rebuilt machine should get
|
||||
a credential is the remaining design question. *(R-193, R-204 item 4)*
|
||||
- **The recovery screen you described has been priced, and it can be built.** A freshly installed
|
||||
machine that finds a sealed package waiting should say so, offer a box for the recovery code, and
|
||||
show what would come back before doing anything. One thing to weigh, deliberately not decided: that
|
||||
screen is reachable by anyone with the household's dashboard password, and the preview reveals
|
||||
backup dates and app names. *(R-193)*
|
||||
- **The orphaned backups on the storage box — you said delete, and it is still owed.** About 1.2 GB
|
||||
across the two demo machines, in set-aside stores nobody can open and nothing prunes. Deliberately not
|
||||
done today: erasing data on the storage endpoint is not something that should ride along with a
|
||||
database change, so it wants its own session. *(R-193)*
|
||||
across the two demo machines, in set-aside stores nobody can open and nothing prunes. It wants its
|
||||
own session rather than riding along with other work. *(R-193)*
|
||||
- *(decided 4 Aug)* You chose **not** to keep a copy of the backup key on the Proxmox host, which
|
||||
makes the customer's own recovery code the only route back from a rebuild. *(R-193)*
|
||||
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
|
||||
session log; nothing suggests anyone else saw it. *(R-132)*
|
||||
- **One small question, not urgent.** The automatic check cannot see which version you have told
|
||||
machines to install, only which ones exist. Closing that needs either a password given to the build
|
||||
server or a check inside the hub itself. *(R-184)*
|
||||
- **Nothing else.**
|
||||
|
||||
## Changed since last update
|
||||
|
||||
- **2026-08-04 (night)** — **The drill PASSED.** We destroyed a machine on purpose and the customer's
|
||||
file came back byte-for-byte identical, using the recovery code you saved. First time the backup
|
||||
story has been proved end to end. It needed four undocumented manual steps to get there, which are
|
||||
now the next piece of work. The machine is up, healthy and re-armed. *(R-201, R-204)*
|
||||
- **2026-08-04 (evening)** — **Fixed the folder-left-out-of-the-backup problem, both halves.** The app
|
||||
and its backup now look in the same directory, and a backup that misses a folder marked essential
|
||||
reports *incomplete* instead of success. Proved by listing the backup's own contents and finding the
|
||||
marked file. The wipe-and-restore proof is unblocked. *(R-203, R-201)*
|
||||
- **2026-08-04 (late)** — Set up the wipe-and-restore proof on the HP machine and **stopped before the
|
||||
wipe**: a folder marked as protected was missing from the off-site backup while the backup reported
|
||||
success. Three things were proved on the way, all firsts: a rebuilt machine's off-site backup now
|
||||
**refuses and says so** instead of quietly starting over; the "start a new store" repair works and
|
||||
keeps the old data aside; and the HP machine's pre-3-August history is gone for good — its key was
|
||||
destroyed four hours before the fix that would have kept it. *(R-203, R-201, R-193)*
|
||||
- **2026-08-04 (evening)** — **Proved the backup key comes back.** A demo machine fetched its own
|
||||
sealed package, opened it with the saved recovery code, and produced a key identical to the one it
|
||||
uses. A wrong code was refused and wrote nothing. The recovery code left no trace anywhere on the
|
||||
machine — searched, with a planted copy first to prove the search worked. *(R-199, R-200)*
|
||||
- **2026-08-04 (earlier)** — **Fixed the worst of it.** The hub now keeps the off-site backup key when a
|
||||
machine re-seals, instead of only the whole-machine one — four lines of database change that stop the
|
||||
recovery ceremony from destroying the thing it exists to protect. A changed backup key now raises an
|
||||
alarm the same day. The daily email that described the opposite of what it measured now states what it
|
||||
saw, and a refusal to auto-repair leaves a record instead of silence. *(R-198, R-197, R-192, R-196)*
|
||||
- **2026-08-04 (later)** — Established what a machine rebuild actually destroys: not the storage
|
||||
password we had been re-issuing, but the key that encrypts the backups. Both demo machines lost
|
||||
access to their off-site history, one of them silently. A decision is now waiting on you. The daily
|
||||
false alarm about David is fixed. *(R-193, R-195)*
|
||||
- **2026-08-04** — Both demo machines proved their on-machine backup restores, on their own,
|
||||
overnight — the copy an ordinary recovery uses, never tested until now. Found while checking: the
|
||||
weekly off-site backup reports failure after a successful upload. *(R-185, R-191)*
|
||||
- **2026-08-03** — Fixed three ways the alarm system misreported itself: a proof of a working backup
|
||||
that could vanish on a restart (seen happening), a release that emailed a failure for a release
|
||||
that worked, and a released binary nobody could rebuild and check. *(R-189, R-188, R-186)*
|
||||
- **One small question, not urgent.** The automatic version check cannot see which version you have
|
||||
told machines to install, only which ones exist. *(R-184)*
|
||||
|
||||
Reference in New Issue
Block a user