a8caa0fdde
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step.
117 lines
8.3 KiB
Markdown
117 lines
8.3 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-22 — the off-site restore now works for all 53 apps, not 13. It is released and
|
|
NOT yet delivered: two steps below are yours.**
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
|
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
|
> If it does not fit, it belongs in the register instead.
|
|
|
|
## Waiting on you
|
|
|
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
|
nothing.*
|
|
|
|
1. **Vouch the golden carrying controller 0.220.2** — Hub → Configuration → Day-0 artifacts.
|
|
**It is already baked, published and round-trip verified**
|
|
(`documentation/tests/golden-0.220.2-2026-08-22/`); only the vouch is left, and only you can do it.
|
|
**It is a THREE-field save:** `golden_version` → **0.220.2**, `agent_version` → **0.130.0**,
|
|
`min_agent` → **0.129.0**. **Then** raise the floor to **0.220.2**, last, in its own save.
|
|
**If you do nothing:** the fleet stays on 0.219.0, so a failed database restore still leaves an app
|
|
broken with an unusable copy — the thing today's release fixes reaches nobody. New machines still
|
|
receive 0.219.0. The build system stays red about it and will mail you on every push.
|
|
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
|
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
|
as it misled one by an hour.
|
|
|
|
## Decided — and what would reopen each
|
|
|
|
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
|
|
asks. *(R-312)*
|
|
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
|
|
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
|
|
abandoned. *(R-313)*
|
|
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
|
|
outside a constructed test. *(R-303)*
|
|
|
|
## What works
|
|
|
|
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
|
|
`demo-hp` runs controller **0.219.0**; the fleet floor is still **0.218.0** (see item 2 above).
|
|
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
|
|
|
|
**The fleet, because two summaries have been misread:** five customer records, three machines.
|
|
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
|
|
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
|
|
**`tester-1` is a record with no machine.**
|
|
|
|
## Shipped
|
|
|
|
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
|
|
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
|
|
already taken a copy of your live database — a good copy — and **nothing in the product could put it
|
|
back.** You were shown a filename. On one of the two database types it was worse: part of the
|
|
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
|
|
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
|
|
app is running. Proven on both database types, byte-identical both times.
|
|
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
|
|
app on a half-written database lets you type into it and makes the damage permanent — and you are
|
|
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
|
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
|
off-site permanently.
|
|
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
|
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
|
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
|
were never given a drive to choose. It was asking one question to answer two. Proven today on
|
|
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
|
|
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
|
|
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
|
|
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
|
|
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
|
|
back, because a restore that mentions only its file count is how a silent loss reads as a success.
|
|
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
|
|
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
|
|
app of 53 was affected**, established with a check first proved able to catch a planted second case.
|
|
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
|
|
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
|
|
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
|
|
healthy.
|
|
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
|
|
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
|
|
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
|
|
The off-site box is back to **17** open connections from **415**.
|
|
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
|
|
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
|
|
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
|
|
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
|
|
can be re-attached after a reinstall** (R-280).
|
|
|
|
## Broken, or knowingly incomplete
|
|
|
|
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not
|
|
the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the
|
|
standard tool — which we never run.
|
|
- **A restore that returns nothing still reports success** (part of R-354's neighbourhood, not fixed
|
|
today) and **verification copies have no delete guard**. Both deliberately left for their own rows.
|
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
|
is **under a year** away on the corrected measurement, not two.
|
|
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
|
|
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
|
|
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
|
|
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
|
|
resolution goes missing quietly.
|
|
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
|
|
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
|
|
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
|
|
fixed twice** (R-327).
|
|
|
|
## Working on next
|
|
|
|
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
|
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|