Files
felhom.eu/STATUS.md
T
admin 8c9f1b798b
gates / gates (push) Successful in 18s
golden 0.219.0 baked, published and round-trip verified (NOT vouched)
Baked in the drill VM per RUNBOOK-manual-build.md 4.0/4.1, carrying controller
v0.219.0 (R-356).

  GOLDEN_VERSION 0.219.0
  GOLDEN_SHA256  67b46f78f8ed9c7b1876265ab1bde9ec6798897898b1836acece9f3864a2aeb6
  656832571 bytes

All five pass markers matched, both negative controls at 0. Verified by ROUND
TRIP - the published object downloaded again and its sha recomputed - not by the
number the script printed.

Both token-leak greps were proved able to convict before their zeros were
believed: planted copy grepped 1, shredded, then the 0 accepted.

Teardown complete: guest 9100 purged, four secret/script files shredded after the
log was copied out, qemu exited, disk reverted to virgin. The revert first
refused while qemu held the image, which is the runbook's own no-holder proof.

NOT vouched - that is a three-field operator save (golden_version 0.219.0,
agent_version 0.130.0, min_agent 0.129.0).
2026-08-22 13:56:33 +02:00

106 lines
7.3 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-08-22 — the off-site restore now works for all 53 apps, not 13. It is released and
NOT yet delivered: two steps below are yours.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
> If it does not fit, it belongs in the register instead.
## Waiting on you
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Vouch the golden carrying controller 0.219.0** — Hub → Configuration → Day-0 artifacts.
**It is baked, published and round-trip verified** (`documentation/tests/golden-0.219.0-2026-08-22/`);
only the vouch is left, and only you can do it. **It is a THREE-field save, not one:**
`golden_version` → **0.219.0**, `agent_version` → **0.130.0**, `min_agent` → **0.129.0**. Moving
`golden_version` alone ships this controller onto an agent older than it declares it needs.
**If you do nothing:** a machine installed today still receives 0.218.0 — the image exists, on the
shelf, undelivered. Reversible: re-select the old values and Save.
2. **Then raise the auto-update floor to 0.219.0 — last, in a separate save.** It acts within seconds.
**If you do nothing:** every existing machine stays on 0.218.0, so the fix below reaches nobody and
40 of 53 apps stay un-restorable on the actual fleet. *(register: R-343's rule)*
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.
## Decided — and what would reopen each
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
asks. *(R-312)*
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
abandoned. *(R-313)*
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
outside a constructed test. *(R-303)*
## What works
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
`demo-hp` runs controller **0.219.0**; the fleet floor is still **0.218.0** (see item 2 above).
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
**The fleet, because two summaries have been misread:** five customer records, three machines.
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
**`tester-1` is a record with no machine.**
## Shipped
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
were never given a drive to choose. It was asking one question to answer two. Proven today on
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
back, because a restore that mentions only its file count is how a silent loss reads as a success.
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
app of 53 was affected**, established with a check first proved able to catch a planted second case.
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
healthy.
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
The off-site box is back to **17** open connections from **415**.
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
can be re-attached after a reinstall** (R-280).
## Broken, or knowingly incomplete
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not
the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the
standard tool — which we never run.
- **A restore that returns nothing still reports success** (part of R-354's neighbourhood, not fixed
today) and **verification copies have no delete guard**. Both deliberately left for their own rows.
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
is **under a year** away on the corrected measurement, not two.
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
resolution goes missing quietly.
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
fixed twice** (R-327).
## Working on next
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).