STATUS + capability map: narrow the end-to-end off-site claim to the leg it was proven on
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
The 2026-08-04 row claimed "a customer's file survives a machine rebuild and comes back — the whole off-site story, end to end". Tonight's drill shows that holds for the declared-userdata leg of a drive-declaring app and for nothing else: the off-site restore has no named-volume leg (R-354), and refuses outright for the 40 apps that declare no data drive (R-356). Since that class keeps ALL its data in named volumes, the end-to-end story is unproven there and disproven for the volume leg generally. The escrow/key half of the row is untouched and still stands. STATUS.md also corrects the fleet pair it still named (0.214.0/0.129.0 -> 0.217.0/0.130.0) and records that demo-hp's off-site had been silent since 9 August.
This commit is contained in:
@@ -1,7 +1,6 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines;
|
||||
it is not yet published).**
|
||||
**Updated 2026-08-21 (night — the backup holds the data; the off-site RESTORE is what loses it).**
|
||||
|
||||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||||
@@ -31,9 +30,14 @@ again.*
|
||||
|
||||
## What works
|
||||
|
||||
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.214.0, agent
|
||||
0.129.0**, delivered by the floor rather than by hand. Off-site is credentialed on `demo-hp` and its
|
||||
repository still opens with the machine's own key. `drill-r50` is reverted to `virgin`, powered off.
|
||||
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.217.0, agent
|
||||
0.130.0**. On `demo-hp` the floor delivered it unaided on 21 August: 0.216.0 → 0.217.0, **17 seconds**
|
||||
from the operator pressing save to the new controller reporting healthy, with no customer action.
|
||||
Off-site is credentialed on `demo-hp`, its repository opens with the machine's own key, and
|
||||
`restic check` over the whole store reports **no errors**. `drill-r50` is reverted to `virgin`, powered
|
||||
off. **`demo-hp` was reinstalled on 21 August and its off-site backup had been silent since 9 August**
|
||||
— the rebuild lost the off-site target, and after that was healed every per-app off-site switch was
|
||||
still off, so the nightly run reported "backup OK" having backed up nothing. Both are on again.
|
||||
|
||||
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
|
||||
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
|
||||
@@ -109,6 +113,29 @@ record with no machine** — created 13 August, no host, no backups, nothing to
|
||||
|
||||
## Broken, or knowingly incomplete
|
||||
|
||||
- **The off-site restore does not give back an app's Docker volume — and says it worked** (R-354).
|
||||
Proven on `demo-hp` on the night of 21 August with files we planted and hashed first. The data IS in
|
||||
the backup: the unit, the off-site snapshot and the verification folder all hold it, byte for byte,
|
||||
accented Hungarian filenames and all. The **full off-site restore then puts back the user files and
|
||||
silently skips the volume**, and tells the customer „5 fájl visszaállítva". For an app that keeps its
|
||||
files on a data drive the lost piece is its settings. **For the 40 of our 53 apps that have no data
|
||||
drive, that volume IS all their data.** The *local* restore does return it correctly — same tar, other
|
||||
code path. **Nothing is lost off-site; the last step is what fails.** *(register: R-354)*
|
||||
- **For those same 40 apps the off-site restore refuses outright, and the reason it gives is untrue**
|
||||
(R-356). It says the app „nincs telepítve" — is not installed — about an app that is installed and
|
||||
running, then tells the customer to reinstall it "to the same place", which those apps give them no
|
||||
way to choose. This is what the OpenGist journey hit on the afternoon of 21 August before falling
|
||||
back to the local restore. *(register: R-356)*
|
||||
- **Paperless's database has never been in its backup** (R-355). Its dump is written every night, is
|
||||
valid, has 72 tables — and lands in a folder named after an app that does not exist, on the wrong
|
||||
disk, where nothing sends it off-site and nothing restores it. Because the same wrong name is used
|
||||
when taking the "undo" copy before a restore, **a restore of Paperless takes no undo at all** and then
|
||||
tells the customer the app has no database. **One app in 53 is affected. It is the document
|
||||
archive.** *(register: R-355)*
|
||||
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the
|
||||
agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the
|
||||
standard tool — which we never run. *(register: R-359)*
|
||||
|
||||
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
|
||||
for a box we actually write to once a week. That volume is what turned a slow internal leak into
|
||||
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
|
||||
|
||||
Reference in New Issue
Block a user