R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s

P1, outcome (i) in one second: replacing the agent on demo-hp released
exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT
stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on
peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another
explanation.

P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is
extrapolated): control +4, fixed +0, with each box making exactly 4
/snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4
leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the
zero is the fix and not a stopped agent.

P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is
precisely the t0 baseline of 2026-08-18 09:51:22Z.

Corrects a claim this session made earlier the same day: the accumulated
descriptors did NOT need an ep0 proxy restart. They were held on both
sides. ep0 was read-only throughout; its PID never changed.

R-344 updated and left OPEN (unpublished is not delivered). R-336
re-scoped -- its old next-step would have fixed nothing while looking like
a failed fix, and it is now a scaling row (~25 req/s at fifty customers).
R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent
restart blanks the reported backup list for ~18 h and the Store comment
calls it unaffected -- blinds no alarm, checked not assumed.
This commit is contained in:
2026-08-20 12:39:32 +02:00
parent 9299f85c4b
commit 57dd62b097
14 changed files with 641 additions and 7 deletions
+17 -5
View File
@@ -1,7 +1,7 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-20 (morning — we looked at who was actually holding the off-site box's connections
open, and it turned out to be us).**
**Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines;
it is not yet published).**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
@@ -124,6 +124,19 @@ record with no machine** — created 13 August, no host, no backups, nothing to
rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed
fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a
deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)*
- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the
idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it
back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one
machine and left the other alone, and in the same hour the untouched one leaked 4 more connections
while the fixed one leaked none — with both doing exactly the same four rounds of work. **The
off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the
built-up connections released themselves when the agents restarted; the off-site box was only ever
read from, never touched. *(register: R-344)*
- **The fix is on the two demo machines by hand and NOT published yet** (R-347). A machine installed
from today's image still gets the old, leaking agent. That was deliberate — publishing it mid-test
would have contaminated the comparison — and the reason has now expired. **It is not urgent:** a new
machine would take the better part of a year to matter, and any agent update clears the build-up.
**Publishing is your call**, and it needs the operator-only artifact screen at the end.
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
installed the newer backup software for the practice, having first read its release notes and found
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
@@ -163,9 +176,8 @@ record with no machine** — created 13 August, no host, no backups, nothing to
## Working on next
**One decision is waiting for you, at the top of `REPORT-spike-ep0-connections.md`:** whether to run the
overnight measurement window on `demo-hp` tonight, and which of two versions of it. Nothing was changed on
any machine in this session — every reading was read-only. After that: the three remaining
**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two
demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
still untriaged against everything since.