# STATUS — what works, what's broken, what's next **Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines; it is not yet published).** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.** > If it does not fit, it belongs in the register instead. ## Waiting on you *Nothing. All three questions that stood here were answered on 12–13 August and have moved to **Decided** below. A decided question left in the deciding list is how a person loses track of what is actually waiting.* ## Decided — and what would reopen each *A decision with no trigger becomes a permanent silence, so each one names what would make us look again.* - **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks — one request from a person who is not us. *(register: R-312)* - **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the only state in existence where a set-aside store is present and cannot be opened, which is the case any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)* - **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed happening outside a constructed test. *(register: R-303)* ## What works Both demo machines are home, healthy and reporting on the approved pair — **controller 0.214.0, agent 0.129.0**, delivered by the floor rather than by hand. Off-site is credentialed on `demo-hp` and its repository still opens with the machine's own key. `drill-r50` is reverted to `virgin`, powered off. **What the fleet actually is, because two summaries have now been misread:** the hub holds **five customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a record with no machine** — created 13 August, no host, no backups, nothing to lose. ## Shipped - **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks only ever watched how full a store was getting, and a failed reading was written to a log nobody reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the weekly backup that happened to fall inside it. After about half an hour of being unable to see a store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back. **Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That second check is written down as the next step. - **The auto-update floor is current again** (R-343). You raised it to today's version this afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*, after the image is vouched, because it acts within seconds. It did: **one machine updated itself nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not paperwork. - **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the 19th", that date is now read by the build system, and a push is refused once it passes. **It is not a reminder service** — it speaks on the next push, not on the day — and that limit is written into the check itself. - **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image had been two releases behind since the 14th — anyone installing would have received a version missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked and published, and **you vouched it**, which was the half that could not be done without you. **The build system is green again for the first time since 14 August**, so the failure mail should stop. The running machines were not touched: this only ever affected *new* installs. - **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out of one internal resource and, while looking perfectly healthy from outside, was accepting no connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a year** on the corrected measurement, not a cure. - **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The second reinstall used to hit our own leftover; it was watched failing on the cycle that actually fails, then watched passing. - **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types the code for an older set of backups, the machine now checks the packages we kept, recognises it, and says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are fine, write to us*. It deliberately promises no restore, because there is no button yet. - **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302). - **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere, on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow „Helyreállítási kód" and cost a real code. - **The hub can see whether a machine's guest still has working networking** (R-319, first reader built against R-264). A machine quietly repairing its own network over and over is now visible instead of being a green tick; a machine that does not report it is drawn as unknown, never as healthy. - **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word from the name we retired last week, and false besides: it restores nothing. Five places, all in the hub; no machine touched. - **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down, then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the morning backup reminder had the same blind spot and is fixed with it. - **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are now checked for a retired name, and the guard has been watched catching one, ignoring an explanation of one, and going quiet again. - **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted. ## Broken, or knowingly incomplete - **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day, for a box we actually write to once a week. That volume is what turned a slow internal leak into last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak itself is untouched.** The honest health check is the resource count climbing, not the absence of an alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years. A deadline, not a comfort. - **The leak is OURS, not Proxmox's, and asking fewer questions would not have fixed it** (R-344, found 2026-08-20). We finally looked at who was on the other end of the stuck connections. Every single one belongs to **our own agent** on the two demo machines — it opens a connection to the off-site box on each 15-minute cycle and never closes it, and neither does the box. The two Proxmox pollers that make 99.5% of the traffic leak **nothing at all**. So the plan recorded under R-336 — turn the question rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)* - **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one machine and left the other alone, and in the same hour the untouched one leaked 4 more connections while the fixed one leaked none — with both doing exactly the same four rounds of work. **The off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the built-up connections released themselves when the agents restarted; the off-site box was only ever read from, never touched. *(register: R-344)* - **PUBLISHED the same day, on your word** (R-347, closed). Agent **0.130.0** is released, and the hub now hands it to any new machine. Both demo machines run the exact published copy. Nothing else on that screen was changed — in particular the controller floor was left alone, and the "minimum agent" setting too, because raising that would have **stopped** machines getting updates rather than helping them. - **Two things the release itself turned up.** (1) The machines were briefly running a *different* build of the same version number — harmless here, but nothing in the system would ever have noticed, because everything compares the version *name*. Now corrected, and filed so it cannot repeat (R-349). (2) **I printed the hub password into my own session log** while confirming the change (R-350). It is not in git and not in any saved file — but it is in the log on this machine. **Changing it is your call**; I can do it without ever showing the new one. Ask and I will. - **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we installed the newer backup software for the practice, having first read its release notes and found **nothing** about the fault we have. The update went cleanly and everything works, but the leak behaves exactly as before, which is the result the release notes predicted. **Two dated checks are booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it. - **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was still showing this morning's failure for four minutes after the copy was safely on the off-site box; `demo-felhom` updated in under a minute. **It corrected itself** and both machines now read correctly, so this is a note to watch, **not something broken** — but during the repair it looked briefly like a second fault, which is the reason it is written down. - **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was not** — its control channel is still bound to the ordinary home network, which is exactly what that change existed to stop. The machine works; the page is wrong, and it misled this session by an hour. Your call which one to correct. - **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a half months and has been silent since 15 July, when its host record was deleted. There is no key, no off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in existence. Whether it stays parked is your call and is deliberately left open. - **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built). Today the honest answer is "your code is right, write to us" — and we can. - **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine; only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight. - **The storage page has its own separate reason for showing an empty list** (R-298), untouched. - **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others are now recorded as deliberately unread, which is honest rather than fixed. - **Two thirds of the standing picture is still unproven, and now you can ask** (R-326). `python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine is how many claims the 9 August review *lowered*, which is a different question. - **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its status may only be raised after the capability map moves first, which is a separate judgement. ## Working on next **One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent); R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292), still untriaged against everything since.