Files
felhom.eu/STATUS.md
T
admin ebfd0967c1
gates / gates (push) Failing after 12s
INCIDENT + registers: ep0's PBS proxy served nobody for 9.5h (R-336..R-338)
Two whole_guest_backup_failed alerts at 04:30 and 04:32 CEST were one
incident, and not on either customer box: ep0's proxmox-backup-proxy was
active, holding its listening socket, and accepting nothing.

Root cause: accept() returning EMFILE. The process held exactly 1024 fds
-- its systemd-default soft RLIMIT_NOFILE -- of which 1016 were sockets
and 547 connections sat in CLOSE-WAIT. The 1024-deep accept backlog had
overflowed (Recv-Q 1025), so every client timed out. It was wedged from
its own loopback too, which is what moved this from a network problem to
a process problem.

Fed by ~85k requests/day (a flat 3,538/hour) against an endpoint written
to weekly, that leak reached the ceiling in 14 days of uptime.

Fix: LimitNOFILE=65536 drop-ins for both PBS units, restart, verified from
both boxes (200 in ~0.1s, felhom-pbs active), then re-drove the missed
backups through the product path -- POST /backup?target=felhom-pbs on each
agent's local API, not a hand-run vzdump.

  demo-felhom  ct/9201/2026-08-18T03:57:43Z  4.10 GB  36.4s
  demo-hp      ct/9201/2026-08-18T03:58:43Z  4.29 GB  41.5s

Both host reports now carry felhom-pbs success=true, so the hub is green on
the evidence rather than on a restart having been performed. No data lost,
no backup skipped: the daily local tier was never affected and the PBS tier
is weekly, so the window cost exactly one attempt.

Evidence copied off ep0 BEFORE the restart, per standing rule 5.

Filed: R-336 (the ~1 req/s poll rate is the real defect; the raised ceiling
is mitigation, not a cure), R-337 (a status endpoint that trailed its own
artifact by minutes then caught up -- WATCHING, downgraded from the defect
I first wrote, because it self-corrected), R-338 (demo-hp is not on the
R-50 island at all and nodes.md says it is; its local API is bound to the
customer LAN).

R-334 updated: still open, now one version wider (controller 0.216.0 vs
golden 0.214.0). golden-currency is the only failing gate and is inherited
-- it reads files this session did not touch -- so this push used
--no-verify, stated per .claude/rules/gates.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN
2026-08-18 06:09:38 +02:00

128 lines
9.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# STATUS — what works, what's broken, what's next
**Updated 2026-08-18 (early — a listening socket that served nobody; off-site backups were down for
9½ hours overnight and are back).**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
> If it does not fit, it belongs in the register instead.
## Waiting on you
*Nothing. All three questions that stood here were answered on 1213 August and have moved to
**Decided** below. A decided question left in the deciding list is how a person loses track of what is
actually waiting.*
## Decided — and what would reopen each
*A decision with no trigger becomes a permanent silence, so each one names what would make us look
again.*
- **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a
support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks —
one request from a person who is not us. *(register: R-312)*
- **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the
only state in existence where a set-aside store is present and cannot be opened, which is the case
any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is
abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)*
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is
unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed
happening outside a constructed test. *(register: R-303)*
## What works
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.214.0, agent
0.129.0**, delivered by the floor rather than by hand. Off-site is credentialed on `demo-hp` and its
repository still opens with the machine's own key. `drill-r50` is reverted to `virgin`, powered off.
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom`
is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a
record with no machine** — created 13 August, no host, no backups, nothing to lose.
## Shipped
- **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their
off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out
of one internal resource and, while looking perfectly healthy from outside, was accepting no
connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups
were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was
skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is
weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys years, not a
cure.
- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The
second reinstall used to hit our own leftover; it was watched failing on the cycle that actually
fails, then watched passing.
- **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types
the code for an older set of backups, the machine now checks the packages we kept, recognises it, and
says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are
fine, write to us*. It deliberately promises no restore, because there is no button yet.
- **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown
banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302).
- **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere,
on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow
„Helyreállítási kód" and cost a real code.
- **The hub can see whether a machine's guest still has working networking** (R-319, first reader built
against R-264). A machine quietly repairing its own network over and over is now visible instead of
being a green tick; a machine that does not report it is drawn as unknown, never as healthy.
- **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an
account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word
from the name we retired last week, and false besides: it restores nothing. Five places, all in the
hub; no machine touched.
- **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down,
then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the
morning backup reminder had the same blind spot and is fixed with it.
- **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are
now checked for a retired name, and the guard has been watched catching one, ignoring an
explanation of one, and going quiet again.
- **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted.
## Broken, or knowingly incomplete
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
for a box we actually write to once a week. That volume is what turned a slow internal leak into
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
itself is untouched.** The honest health check is the resource count climbing, not the absence of an
alarm.
- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was
still showing this morning's failure for four minutes after the copy was safely on the off-site box;
`demo-felhom` updated in under a minute. **It corrected itself** and both machines now read
correctly, so this is a note to watch, **not something broken** — but during the repair it looked
briefly like a second fault, which is the reason it is written down.
- **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both
machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was
not** — its control channel is still bound to the ordinary home network, which is exactly what that
change existed to stop. The machine works; the page is wrong, and it misled this session by an hour.
Your call which one to correct.
- **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine
belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a
half months and has been silent since 15 July, when its host record was deleted. There is no key, no
off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the
visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in
existence. Whether it stays parked is your call and is deliberately left open.
- **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built).
Today the honest answer is "your code is right, write to us" — and we can.
- **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine;
only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight.
- **The storage page has its own separate reason for showing an empty list** (R-298), untouched.
- **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent
update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others
are now recorded as deliberately unread, which is honest rather than fixed.
- **Two thirds of the standing picture is still unproven, and now you can ask** (R-326).
`python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of
those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine
is how many claims the 9 August review *lowered*, which is a different question.
- **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its
status may only be raised after the capability map moves first, which is a separate judgement.
## Working on next
`demo-hp` is yours this evening — **this session did not touch it**. After that: the three remaining
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
still untriaged against everything since.