57dd62b097
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed.
184 lines
14 KiB
Markdown
184 lines
14 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
||
|
||
**Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines;
|
||
it is not yet published).**
|
||
|
||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||
> If it does not fit, it belongs in the register instead.
|
||
|
||
## Waiting on you
|
||
|
||
*Nothing. All three questions that stood here were answered on 12–13 August and have moved to
|
||
**Decided** below. A decided question left in the deciding list is how a person loses track of what is
|
||
actually waiting.*
|
||
|
||
## Decided — and what would reopen each
|
||
|
||
*A decision with no trigger becomes a permanent silence, so each one names what would make us look
|
||
again.*
|
||
|
||
- **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a
|
||
support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks —
|
||
one request from a person who is not us. *(register: R-312)*
|
||
- **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the
|
||
only state in existence where a set-aside store is present and cannot be opened, which is the case
|
||
any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is
|
||
abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)*
|
||
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is
|
||
unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed
|
||
happening outside a constructed test. *(register: R-303)*
|
||
|
||
## What works
|
||
|
||
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.214.0, agent
|
||
0.129.0**, delivered by the floor rather than by hand. Off-site is credentialed on `demo-hp` and its
|
||
repository still opens with the machine's own key. `drill-r50` is reverted to `virgin`, powered off.
|
||
|
||
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
|
||
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
|
||
disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom`
|
||
is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a
|
||
record with no machine** — created 13 August, no host, no backups, nothing to lose.
|
||
|
||
## Shipped
|
||
|
||
- **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a
|
||
completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks
|
||
only ever watched how full a store was getting, and a failed reading was written to a log nobody
|
||
reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the
|
||
weekly backup that happened to fall inside it. After about half an hour of being unable to see a
|
||
store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back.
|
||
**Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have
|
||
caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That
|
||
second check is written down as the next step.
|
||
|
||
- **The auto-update floor is current again** (R-343). You raised it to today's version this
|
||
afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*,
|
||
after the image is vouched, because it acts within seconds. It did: **one machine updated itself
|
||
nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not
|
||
paperwork.
|
||
- **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the
|
||
19th", that date is now read by the build system, and a push is refused once it passes. **It is not
|
||
a reminder service** — it speaks on the next push, not on the day — and that limit is written into
|
||
the check itself.
|
||
|
||
- **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image
|
||
had been two releases behind since the 14th — anyone installing would have received a version
|
||
missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked
|
||
and published, and **you vouched it**, which was the half that could not be done without you. **The
|
||
build system is green again for the first time since 14 August**, so the failure mail should stop.
|
||
The running machines were not touched: this only ever affected *new* installs.
|
||
|
||
- **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their
|
||
off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out
|
||
of one internal resource and, while looking perfectly healthy from outside, was accepting no
|
||
connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups
|
||
were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was
|
||
skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is
|
||
weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question
|
||
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a
|
||
year** on the corrected measurement, not a cure.
|
||
|
||
- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The
|
||
second reinstall used to hit our own leftover; it was watched failing on the cycle that actually
|
||
fails, then watched passing.
|
||
- **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types
|
||
the code for an older set of backups, the machine now checks the packages we kept, recognises it, and
|
||
says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are
|
||
fine, write to us*. It deliberately promises no restore, because there is no button yet.
|
||
- **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown
|
||
banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302).
|
||
- **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere,
|
||
on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow
|
||
„Helyreállítási kód" and cost a real code.
|
||
- **The hub can see whether a machine's guest still has working networking** (R-319, first reader built
|
||
against R-264). A machine quietly repairing its own network over and over is now visible instead of
|
||
being a green tick; a machine that does not report it is drawn as unknown, never as healthy.
|
||
- **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an
|
||
account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word
|
||
from the name we retired last week, and false besides: it restores nothing. Five places, all in the
|
||
hub; no machine touched.
|
||
- **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down,
|
||
then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the
|
||
morning backup reminder had the same blind spot and is fixed with it.
|
||
- **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are
|
||
now checked for a retired name, and the guard has been watched catching one, ignoring an
|
||
explanation of one, and going quiet again.
|
||
- **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted.
|
||
|
||
## Broken, or knowingly incomplete
|
||
|
||
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
|
||
for a box we actually write to once a week. That volume is what turned a slow internal leak into
|
||
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
|
||
itself is untouched.** The honest health check is the resource count climbing, not the absence of an
|
||
alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too
|
||
optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years.
|
||
A deadline, not a comfort.
|
||
- **The leak is OURS, not Proxmox's, and asking fewer questions would not have fixed it** (R-344, found
|
||
2026-08-20). We finally looked at who was on the other end of the stuck connections. Every single one
|
||
belongs to **our own agent** on the two demo machines — it opens a connection to the off-site box on
|
||
each 15-minute cycle and never closes it, and neither does the box. The two Proxmox pollers that make
|
||
99.5% of the traffic leak **nothing at all**. So the plan recorded under R-336 — turn the question
|
||
rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed
|
||
fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a
|
||
deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)*
|
||
- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the
|
||
idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it
|
||
back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one
|
||
machine and left the other alone, and in the same hour the untouched one leaked 4 more connections
|
||
while the fixed one leaked none — with both doing exactly the same four rounds of work. **The
|
||
off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the
|
||
built-up connections released themselves when the agents restarted; the off-site box was only ever
|
||
read from, never touched. *(register: R-344)*
|
||
- **The fix is on the two demo machines by hand and NOT published yet** (R-347). A machine installed
|
||
from today's image still gets the old, leaking agent. That was deliberate — publishing it mid-test
|
||
would have contaminated the comparison — and the reason has now expired. **It is not urgent:** a new
|
||
machine would take the better part of a year to matter, and any agent update clears the build-up.
|
||
**Publishing is your call**, and it needs the operator-only artifact screen at the end.
|
||
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
|
||
installed the newer backup software for the practice, having first read its release notes and found
|
||
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
|
||
behaves exactly as before, which is the result the release notes predicted. **Two dated checks are
|
||
booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it.
|
||
- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was
|
||
still showing this morning's failure for four minutes after the copy was safely on the off-site box;
|
||
`demo-felhom` updated in under a minute. **It corrected itself** and both machines now read
|
||
correctly, so this is a note to watch, **not something broken** — but during the repair it looked
|
||
briefly like a second fault, which is the reason it is written down.
|
||
- **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both
|
||
machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was
|
||
not** — its control channel is still bound to the ordinary home network, which is exactly what that
|
||
change existed to stop. The machine works; the page is wrong, and it misled this session by an hour.
|
||
Your call which one to correct.
|
||
|
||
- **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine
|
||
belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a
|
||
half months and has been silent since 15 July, when its host record was deleted. There is no key, no
|
||
off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the
|
||
visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in
|
||
existence. Whether it stays parked is your call and is deliberately left open.
|
||
- **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built).
|
||
Today the honest answer is "your code is right, write to us" — and we can.
|
||
- **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine;
|
||
only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight.
|
||
- **The storage page has its own separate reason for showing an empty list** (R-298), untouched.
|
||
- **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent
|
||
update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others
|
||
are now recorded as deliberately unread, which is honest rather than fixed.
|
||
- **Two thirds of the standing picture is still unproven, and now you can ask** (R-326).
|
||
`python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of
|
||
those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine
|
||
is how many claims the 9 August review *lowered*, which is a different question.
|
||
- **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its
|
||
status may only be raised after the capability map moves first, which is a separate judgement.
|
||
|
||
## Working on next
|
||
|
||
**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two
|
||
demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining
|
||
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
|
||
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
|
||
still untriaged against everything since.
|