Files
felhom.eu/STATUS.md
T
admin 8c9f1b798b
gates / gates (push) Successful in 18s
golden 0.219.0 baked, published and round-trip verified (NOT vouched)
Baked in the drill VM per RUNBOOK-manual-build.md 4.0/4.1, carrying controller
v0.219.0 (R-356).

  GOLDEN_VERSION 0.219.0
  GOLDEN_SHA256  67b46f78f8ed9c7b1876265ab1bde9ec6798897898b1836acece9f3864a2aeb6
  656832571 bytes

All five pass markers matched, both negative controls at 0. Verified by ROUND
TRIP - the published object downloaded again and its sha recomputed - not by the
number the script printed.

Both token-leak greps were proved able to convict before their zeros were
believed: planted copy grepped 1, shredded, then the 0 accepted.

Teardown complete: guest 9100 purged, four secret/script files shredded after the
log was copied out, qemu exited, disk reverted to virgin. The revert first
refused while qemu held the image, which is the runbook's own no-holder proof.

NOT vouched - that is a three-field operator save (golden_version 0.219.0,
agent_version 0.130.0, min_agent 0.129.0).
2026-08-22 13:56:33 +02:00

7.3 KiB

STATUS — what works, what's broken, what's next

Updated 2026-08-22 — the off-site restore now works for all 53 apps, not 13. It is released and NOT yet delivered: two steps below are yours.

A view, not a source. documentation/backlog/OPEN-ITEMS.md is the authority; this page restates part of it in plain words, and nothing may exist only here. Items, not paragraphs. One screen. If it does not fit, it belongs in the register instead.

Waiting on you

This section is allowed to be longer than one screen, and each item says what happens if you do nothing.

  1. Vouch the golden carrying controller 0.219.0 — Hub → Configuration → Day-0 artifacts. It is baked, published and round-trip verified (documentation/tests/golden-0.219.0-2026-08-22/); only the vouch is left, and only you can do it. It is a THREE-field save, not one: golden_version → 0.219.0, agent_version → 0.130.0, min_agent → 0.129.0. Moving golden_version alone ships this controller onto an agent older than it declares it needs. If you do nothing: a machine installed today still receives 0.218.0 — the image exists, on the shelf, undelivered. Reversible: re-select the old values and Save.
  2. Then raise the auto-update floor to 0.219.0 — last, in a separate save. It acts within seconds. If you do nothing: every existing machine stays on 0.218.0, so the fix below reaches nobody and 40 of 53 apps stay un-restorable on the actual fleet. (register: R-343's rule)
  3. Whether to change the hub password (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. If you do nothing: it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
  4. demo-hp's network setup does not match our own notes (R-338) — the machine works, the page is wrong, or the other way round. If you do nothing: the page keeps misleading the next session, as it misled one by an hour.

Decided — and what would reopen each

  • Getting old backups back yourself: NOT BUILT, deliberately. Reopens if: a real customer asks. (R-312)
  • The unopenable old copy on demo-felhom: KEPT as a test fixture — the only state in existence where a set-aside store is present and cannot be opened. Delete when: that work ships or is abandoned. (R-313)
  • A machine in two kinds of trouble says both things: LEFT AS IT IS. Reopens if: observed outside a constructed test. (R-303)

What works

Both demo machines are home, healthy and reporting — agent 0.130.0 published and running on both. demo-hp runs controller 0.219.0; the fleet floor is still 0.218.0 (see item 2 above). Off-site is credentialed on demo-hp and its store opens with the machine's own key.

The fleet, because two summaries have been misread: five customer records, three machines. demo-felhom and demo-hp are ours and disposable; drill-r50 is a nested drill VM, reverted and off. peti-felhom is a real machine we have not heard from since 15 July and has no host record. tester-1 is a record with no machine.

Shipped

  • The off-site restore now works for the other 40 apps (R-356, controller 0.219.0, proven on demo-hp). It used to refuse before starting, tell the customer a running app „nincs telepítve", and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they were never given a drive to choose. It was asking one question to answer two. Proven today on privatebin: data planted through the app itself, backed up, deleted, restored — all 15 files back byte for byte, Hungarian accented names included, message „0 fájl és 1 adatkötet visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
  • The off-site restore gives an app's data back at all (R-354, controller 0.218.0). It used to say „0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came back, because a restore that mentions only its file count is how a silent loss reads as a success.
  • Paperless's database is in the backup, and restoring it takes an undo copy first (R-355, controller 0.218.0). The dump was landing in a folder named after an app that does not exist. One app of 53 was affected, established with a check first proved able to catch a planted second case.
  • The system tells you when it cannot see the off-site copies (R-339) — a mail after ~30 minutes, hourly while it lasts, one all-clear. Caveat: it watches whether the machine answers, so it would not have caught the 18 August fault, where one service was wedged and the machine stayed healthy.
  • The connection leak was ours and is fixed (R-344). Our agent opened a connection to the off-site box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one. The off-site box is back to 17 open connections from 415.
  • A dated check can no longer be quietly missed (R-341) — but it speaks on the next push, not on the day. A machine we tell to be quiet is no longer reported as dead (R-321). One name per secret (R-295, R-323). The hub's own words are under a guard (R-324). Removal reverses the installation (R-316). A correct recovery code is no longer called wrong (R-311). The drive can be re-attached after a reinstall (R-280).

Broken, or knowingly incomplete

  • Nothing ever checks that the off-site store is still readable (R-359). Not the controller, not the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the standard tool — which we never run.
  • A restore that returns nothing still reports success (part of R-354's neighbourhood, not fixed today) and verification copies have no delete guard. Both deliberately left for their own rows.
  • We ask the off-site box a question about once a second (R-336) — ~85,000 a day for a box we write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling is under a year away on the corrected measurement, not two.
  • Peti's machine has no recovery route at all. A real machine belonging to a real person, silent since 15 July, no key, no off-site copy, no local backup. If that drive fails, everything on it is lost. First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
  • The agent picks dnsmasq by looking at a file another package owns (R-317) — one line; LAN name resolution goes missing quietly.
  • Three facts the machines send still have no reader (R-264); the storage page has its own reason for an empty list (R-298); two thirds of the standing picture is unproven (R-326: 23 of 55 claims walked — python3 scripts/unproven.py); the picture still describes one defect we fixed twice (R-327).

Working on next

The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).