STATUS + capability map: narrow the end-to-end off-site claim to the leg it was proven on
gates / gates (push) Successful in 16s

The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back
— the whole off-site story, end to end". Tonight's drill shows that holds for the
declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore
has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data
drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story
is unproven there and disproven for the volume leg generally. The escrow/key half of the
row is untouched and still stands.

STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0)
and records that demo-hp's off-site had been silent since 9 August.
This commit is contained in:
2026-08-21 23:34:21 +02:00
parent f5a4fceeeb
commit d895d9f7dd
3 changed files with 36 additions and 9 deletions
+32 -5
View File
@@ -1,7 +1,6 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines;
it is not yet published).**
**Updated 2026-08-21 (night — the backup holds the data; the off-site RESTORE is what loses it).**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
@@ -31,9 +30,14 @@ again.*
## What works
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.214.0, agent
0.129.0**, delivered by the floor rather than by hand. Off-site is credentialed on `demo-hp` and its
repository still opens with the machine's own key. `drill-r50` is reverted to `virgin`, powered off.
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.217.0, agent
0.130.0**. On `demo-hp` the floor delivered it unaided on 21 August: 0.216.0 → 0.217.0, **17 seconds**
from the operator pressing save to the new controller reporting healthy, with no customer action.
Off-site is credentialed on `demo-hp`, its repository opens with the machine's own key, and
`restic check` over the whole store reports **no errors**. `drill-r50` is reverted to `virgin`, powered
off. **`demo-hp` was reinstalled on 21 August and its off-site backup had been silent since 9 August**
— the rebuild lost the off-site target, and after that was healed every per-app off-site switch was
still off, so the nightly run reported "backup OK" having backed up nothing. Both are on again.
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
@@ -109,6 +113,29 @@ record with no machine** — created 13 August, no host, no backups, nothing to
## Broken, or knowingly incomplete
- **The off-site restore does not give back an app's Docker volume — and says it worked** (R-354).
Proven on `demo-hp` on the night of 21 August with files we planted and hashed first. The data IS in
the backup: the unit, the off-site snapshot and the verification folder all hold it, byte for byte,
accented Hungarian filenames and all. The **full off-site restore then puts back the user files and
silently skips the volume**, and tells the customer „5 fájl visszaállítva". For an app that keeps its
files on a data drive the lost piece is its settings. **For the 40 of our 53 apps that have no data
drive, that volume IS all their data.** The *local* restore does return it correctly — same tar, other
code path. **Nothing is lost off-site; the last step is what fails.** *(register: R-354)*
- **For those same 40 apps the off-site restore refuses outright, and the reason it gives is untrue**
(R-356). It says the app „nincs telepítve" — is not installed — about an app that is installed and
running, then tells the customer to reinstall it "to the same place", which those apps give them no
way to choose. This is what the OpenGist journey hit on the afternoon of 21 August before falling
back to the local restore. *(register: R-356)*
- **Paperless's database has never been in its backup** (R-355). Its dump is written every night, is
valid, has 72 tables — and lands in a folder named after an app that does not exist, on the wrong
disk, where nothing sends it off-site and nothing restores it. Because the same wrong name is used
when taking the "undo" copy before a restore, **a restore of Paperless takes no undo at all** and then
tells the customer the app has no database. **One app in 53 is affected. It is the document
archive.** *(register: R-355)*
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the
agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the
standard tool — which we never run. *(register: R-359)*
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
for a box we actually write to once a week. That volume is what turned a slow internal leak into
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak