R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed.
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-08-20 (morning — we looked at who was actually holding the off-site box's connections
|
||||
open, and it turned out to be us).**
|
||||
**Updated 2026-08-20 (midday — the leak was ours, and it is fixed and proven on both demo machines;
|
||||
it is not yet published).**
|
||||
|
||||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||||
@@ -124,6 +124,19 @@ record with no machine** — created 13 August, no host, no backups, nothing to
|
||||
rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed
|
||||
fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a
|
||||
deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)*
|
||||
- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the
|
||||
idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it
|
||||
back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one
|
||||
machine and left the other alone, and in the same hour the untouched one leaked 4 more connections
|
||||
while the fixed one leaked none — with both doing exactly the same four rounds of work. **The
|
||||
off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the
|
||||
built-up connections released themselves when the agents restarted; the off-site box was only ever
|
||||
read from, never touched. *(register: R-344)*
|
||||
- **The fix is on the two demo machines by hand and NOT published yet** (R-347). A machine installed
|
||||
from today's image still gets the old, leaking agent. That was deliberate — publishing it mid-test
|
||||
would have contaminated the comparison — and the reason has now expired. **It is not urgent:** a new
|
||||
machine would take the better part of a year to matter, and any agent update clears the build-up.
|
||||
**Publishing is your call**, and it needs the operator-only artifact screen at the end.
|
||||
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
|
||||
installed the newer backup software for the practice, having first read its release notes and found
|
||||
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
|
||||
@@ -163,9 +176,8 @@ record with no machine** — created 13 August, no host, no backups, nothing to
|
||||
|
||||
## Working on next
|
||||
|
||||
**One decision is waiting for you, at the top of `REPORT-spike-ep0-connections.md`:** whether to run the
|
||||
overnight measurement window on `demo-hp` tonight, and which of two versions of it. Nothing was changed on
|
||||
any machine in this session — every reading was read-only. After that: the three remaining
|
||||
**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two
|
||||
demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining
|
||||
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
|
||||
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
|
||||
still untriaged against everything since.
|
||||
|
||||
Reference in New Issue
Block a user