# STATUS — what works, what's broken, what's next **Updated 2026-08-22 — both of last night's worst findings are fixed and proven on the machine.** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.** > If it does not fit, it belongs in the register instead. ## Waiting on you *Nothing. All three questions that stood here were answered on 12–13 August and have moved to **Decided** below. A decided question left in the deciding list is how a person loses track of what is actually waiting.* ## Decided — and what would reopen each *A decision with no trigger becomes a permanent silence, so each one names what would make us look again.* - **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks — one request from a person who is not us. *(register: R-312)* - **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the only state in existence where a set-aside store is present and cannot be opened, which is the case any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)* - **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed happening outside a constructed test. *(register: R-303)* ## What works Both demo machines are home, healthy and reporting on the approved pair — **controller 0.217.0, agent 0.130.0**. On `demo-hp` the floor delivered it unaided on 21 August: 0.216.0 → 0.217.0, **17 seconds** from the operator pressing save to the new controller reporting healthy, with no customer action. Off-site is credentialed on `demo-hp`, its repository opens with the machine's own key, and `restic check` over the whole store reports **no errors**. `drill-r50` is reverted to `virgin`, powered off. **`demo-hp` was reinstalled on 21 August and its off-site backup had been silent since 9 August** — the rebuild lost the off-site target, and after that was healed every per-app off-site switch was still off, so the nightly run reported "backup OK" having backed up nothing. Both are on again. **What the fleet actually is, because two summaries have now been misread:** the hub holds **five customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a record with no machine** — created 13 August, no host, no backups, nothing to lose. ## Shipped - **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks only ever watched how full a store was getting, and a failed reading was written to a log nobody reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the weekly backup that happened to fall inside it. After about half an hour of being unable to see a store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back. **Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That second check is written down as the next step. - **The auto-update floor is current again** (R-343). You raised it to today's version this afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*, after the image is vouched, because it acts within seconds. It did: **one machine updated itself nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not paperwork. - **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the 19th", that date is now read by the build system, and a push is refused once it passes. **It is not a reminder service** — it speaks on the next push, not on the day — and that limit is written into the check itself. - **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image had been two releases behind since the 14th — anyone installing would have received a version missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked and published, and **you vouched it**, which was the half that could not be done without you. **The build system is green again for the first time since 14 August**, so the failure mail should stop. The running machines were not touched: this only ever affected *new* installs. - **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out of one internal resource and, while looking perfectly healthy from outside, was accepting no connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a year** on the corrected measurement, not a cure. - **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The second reinstall used to hit our own leftover; it was watched failing on the cycle that actually fails, then watched passing. - **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types the code for an older set of backups, the machine now checks the packages we kept, recognises it, and says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are fine, write to us*. It deliberately promises no restore, because there is no button yet. - **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302). - **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere, on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow „Helyreállítási kód" and cost a real code. - **The hub can see whether a machine's guest still has working networking** (R-319, first reader built against R-264). A machine quietly repairing its own network over and over is now visible instead of being a green tick; a machine that does not report it is drawn as unknown, never as healthy. - **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word from the name we retired last week, and false besides: it restores nothing. Five places, all in the hub; no machine touched. - **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down, then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the morning backup reminder had the same blind spot and is fixed with it. - **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are now checked for a retired name, and the guard has been watched catching one, ignoring an explanation of one, and going quiet again. - **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted. ## Broken, or knowingly incomplete - **FIXED and proven on the machine: the off-site restore gives the app's data back** (R-354, controller 0.218.0). The same planted files, the same steps, both runs on `demo-hp`: on the old build the restore said „0 fájl visszaállítva", reported success, and the folder was simply not there. On the new one it says „**0 fájl és 1 adatkötet visszaállítva**" and all five files come back **byte for byte**, Hungarian accented names included. The message now names what came back, because a restore that mentions only its file count is how a silent loss reads as a success. *(register: R-354, CLOSED)* - **FIXED and proven on the machine: Paperless's database is in the backup, and a restore of it now takes an undo copy first** (R-355, controller 0.218.0). The dump was going into a folder named after an app that does not exist, so nothing collected it — and because the same wrong name was used when looking for the live database, a restore took **no undo copy at all**. Now: the dump is in the app's own backup and in the off-site copy for the first time; the restore said „**0 fájl és 3 adatkötet és az adatbázis visszaállítva**"; and with the undo deliberately made impossible the restore **refused and did not even stop the app**. We no longer guess which app a database belongs to — Docker already tells us. **One app of 53 was affected**, established with a check we first proved could catch a planted second case. *(register: R-355, CLOSED)* - **STILL BROKEN, and it is now the one that matters most: 40 of our 53 apps still cannot use the off-site restore at all** (R-356). It refuses before it starts, says a running app „nincs telepítve" — is not installed — and tells the customer to reinstall it "to the same place", which those apps give them no way to choose. **Those are exactly the apps whose entire data is the thing R-354 just fixed**, so today's fix cannot reach them until this one is done. *(register: R-356)* - **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the standard tool — which we never run. *(register: R-359)* - **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day, for a box we actually write to once a week. That volume is what turned a slow internal leak into last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak itself is untouched.** The honest health check is the resource count climbing, not the absence of an alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years. A deadline, not a comfort. - **The leak is OURS, not Proxmox's, and asking fewer questions would not have fixed it** (R-344, found 2026-08-20). We finally looked at who was on the other end of the stuck connections. Every single one belongs to **our own agent** on the two demo machines — it opens a connection to the off-site box on each 15-minute cycle and never closes it, and neither does the box. The two Proxmox pollers that make 99.5% of the traffic leak **nothing at all**. So the plan recorded under R-336 — turn the question rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)* - **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one machine and left the other alone, and in the same hour the untouched one leaked 4 more connections while the fixed one leaked none — with both doing exactly the same four rounds of work. **The off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the built-up connections released themselves when the agents restarted; the off-site box was only ever read from, never touched. *(register: R-344)* - **PUBLISHED the same day, on your word** (R-347, closed). Agent **0.130.0** is released, and the hub now hands it to any new machine. Both demo machines run the exact published copy. Nothing else on that screen was changed — in particular the controller floor was left alone, and the "minimum agent" setting too, because raising that would have **stopped** machines getting updates rather than helping them. - **Two things the release itself turned up.** (1) The machines were briefly running a *different* build of the same version number — harmless here, but nothing in the system would ever have noticed, because everything compares the version *name*. Now corrected, and filed so it cannot repeat (R-349). (2) **I printed the hub password into my own session log** while confirming the change (R-350). It is not in git and not in any saved file — but it is in the log on this machine. **Changing it is your call**; I can do it without ever showing the new one. Ask and I will. - **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we installed the newer backup software for the practice, having first read its release notes and found **nothing** about the fault we have. The update went cleanly and everything works, but the leak behaves exactly as before, which is the result the release notes predicted. **Two dated checks are booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it. - **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was still showing this morning's failure for four minutes after the copy was safely on the off-site box; `demo-felhom` updated in under a minute. **It corrected itself** and both machines now read correctly, so this is a note to watch, **not something broken** — but during the repair it looked briefly like a second fault, which is the reason it is written down. - **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was not** — its control channel is still bound to the ordinary home network, which is exactly what that change existed to stop. The machine works; the page is wrong, and it misled this session by an hour. Your call which one to correct. - **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a half months and has been silent since 15 July, when its host record was deleted. There is no key, no off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in existence. Whether it stays parked is your call and is deliberately left open. - **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built). Today the honest answer is "your code is right, write to us" — and we can. - **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine; only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight. - **The storage page has its own separate reason for showing an empty list** (R-298), untouched. - **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others are now recorded as deliberately unread, which is honest rather than fixed. - **Two thirds of the standing picture is still unproven, and now you can ask** (R-326). `python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine is how many claims the 9 August review *lowered*, which is a different question. - **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its status may only be raised after the capability map moves first, which is a separate judgement. ## Working on next **One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent); R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292), still untriaged against everything since.