Read-only diagnostic; nothing restarted, reconfigured or deployed. CAUSE (H2, supported): the island migration rewrote bootstrap.json (169.254.253.1:8443) but the controller reads controller.yaml, which still holds the pre-island LAN address on both boxes -- mtimes 2026-07-21 and 2026-07-23, i.e. days BEFORE the migration. ensureLocalAPI (bootstrap.go:254) returns early when local_api.endpoint is non-empty, so it fills a MISSING endpoint and never reconciles a STALE one. The agent no longer binds the LAN address, hence "connection refused". H1 refuted: every island leg is healthy on both hosts -- agent active and bound 169.254.253.1:8443, vmbr9 up, guest eth1 up, TCP from the guest to the island is OPEN. The island works; nothing is using it. H3 refuted: hub-persisted events put the first alarm at 2026-07-25 12:44/12:46, ~17.5h before the v0.172.0 deploy. Today's re-fires are restart artifacts -- channel state is in-memory by design. H4 refuted: felhom-agent is active on both hosts. Also: the logged "up->down" is an ARTIFACT -- the debounce branch seeds an unseeded state to "up", so the channel was born-down and has never been up. Symptom B confirmed: EffectiveProtected gates cloudflared on a token predicate but samba on smb.Enabled alone, so sharing-without-password reports FAIL for a deliberately-undeployed stack. The function's own doc comment asserts the parity that is broken. Caused by leg 4 of the R-75 task. Blast radius: zero emails so far, but demo-hp flipped ok->fail and the health_critical path has reached the CUSTOMER channel historically. The English issue strings are whole-surface, not a one-off. Fixes described, none applied.
21 KiB
DIAG — agent-channel banner + felhom-samba false alarm (2026-07-26)
Read-only diagnostic. Nothing was restarted, reloaded, reconfigured or deployed. No
systemctl action, no ifreload/ifup, no config edit, no runbook leg, no sharing change, no
container created or removed. Every probe was a read, a log tail, or a kernel route lookup
(ip route get, which sends no packets). Two TCP connect attempts were made from guest shells —
explicitly sanctioned by the brief as non-mutating.
Verdict up front, and it is unambiguous:
The island migration rewrote
bootstrap.jsonbut the controller readscontroller.yaml, andensureLocalAPIonly fills alocal_api.endpointthat is missing — never one that is stale. Both controllers are still dialling the pre-island LAN address, where the agent no longer listens. H2 SUPPORTED. H1, H3, H4 all REFUTED by direct observation.
Symptom B is confirmed, and is a side-effect of leg 4 of the R-75 task (sharing enabled on demo-hp without a household password) — see §5, including the honest note that this diagnostic's own author caused it.
1. D0 — access inventory
| Target | Reachable | Path | User | sudo |
|---|---|---|---|---|
| demo-felhom PVE host | YES | ssh felhom-pve (Tailscale alias) |
root (uid 0) |
n/a — already root |
| demo-felhom guest 9201 | YES | pct exec 9201 from the host |
root in guest | n/a |
| demo-hp PVE host | YES | sshpass … root@demo-hp (TS 100.76.96.79) |
root (uid 0) |
n/a |
| demo-hp guest 9201 | YES | pct exec 9201 from the host |
root in guest | n/a |
How demo-hp was reached, stated plainly: it has no baked operator key. I used the documented
G1 break-glass — the root password vaulted in the hub host_recovery/demo-hp-bb76ea, per
felhom-controller/CLAUDE.md:133 and operations/nodes.md. That is the sanctioned and only access
path to that box, not an escalation. The hub DB copy and the plaintext password were shred -u'd
immediately after use, and no credential appears in this document.
No access gaps. Everything §D1–§D6 asked for was collectable except one item, recorded in §7.
Exit-code discipline. Most probes ran through ssh → pct exec → bash -c pipelines where $? is
the last stage's status, so I read stderr and content, not rc. Three probes failed visibly and
were re-run: two quoting failures (fixed by base64-encoding the probe script — the
ssh → pct exec → bash -c escaping chain mangles (, ' and $), and one wrong SQLite column
name. agent.json reported "unreadable" — recorded in §7 rather than glossed.
2. D1/D2/D3 — Symptom A evidence table
| demo-felhom | demo-hp | |
|---|---|---|
| Reason | unreachable |
unreachable |
| Which branch | checker.go:94-96 (connection refused) — NOT the default unknown case |
same |
| Raw error | agentapi: GET /storage: Get "https://192.168.0.162:8443/storage": dial tcp 192.168.0.162:8443: connect: connection refused |
agentapi: GET /storage: Get "https://192.168.0.87:8443/storage": dial tcp 192.168.0.87:8443: connect: connection refused |
| Transition logged | up->down:unreachable |
up->down:unreachable (preceded by transient down (unreachable, 1/2)) |
| Actually born-down? | YES — see §2.1 | YES |
| Ever recovered? | No agent_channel_recovered since 2026-07-20 |
No, ever |
| First seen (hub, persisted) | 2026-07-25 12:46:42 UTC | 2026-07-25 12:44:19 UTC |
controller.yaml endpoint |
192.168.0.162:8443 (old LAN) |
192.168.0.87:8443 (old LAN) |
controller.yaml mtime |
2026-07-21 08:29:31 | 2026-07-23 09:53:39 |
bootstrap.json endpoint |
169.254.253.1:8443 (island) |
169.254.253.1:8443 (island) |
bootstrap.json mtime |
2026-07-25 12:44:41 | 2026-07-25 12:42:17 |
| Agree? | NO | NO |
| token / fingerprint | both present; fingerprint 64 chars | both present; fingerprint 64 chars |
| Agent listening | 169.254.253.1:8443 only |
169.254.253.1:8443 only |
felhom-agent |
active, since 2026-07-25 14:44:37 CEST |
active, since 2026-07-25 14:42:13 CEST |
vmbr9 |
169.254.253.1/30, up |
169.254.253.1/30, up |
vmbr0 |
192.168.0.162/24, normal |
192.168.0.87/24, normal |
| Persistent vmbr9 stanza | present in /etc/network/interfaces |
present |
Guest net1 |
eth1,bridge=vmbr9,ip=169.254.253.2/30 |
same |
Guest eth1 |
169.254.253.2/30, up |
169.254.253.2/30, up |
| Host uptime | up since 2026-07-22 19:13 (no reboot since migration) | up since 2026-07-23 20:30 (no reboot since migration) |
TCP from guest → island 169.254.253.1:8443 |
OPEN | OPEN |
TCP from guest → old LAN :8443 |
REFUSED | REFUSED |
| TCP from container | not collectable — the image has no nc/wget/ip; answered by route simulation instead (§2.2) |
same |
(Tokens and fingerprints were redacted at the point of collection — the probe script rewrote them to
<REDACTED> before the value ever left the box.)
2.1 up->down is an ARTIFACT — the channel was never up
checker.go seeds an unseeded state to "up" inside the debounce branch:
if cls.debounce && c.consecutiveDown < debounceThreshold {
if c.state == "" { c.state = "up" } // <-- unseeded assumed up until confirmed
So the sequence is: unseeded → probe 1 fails → state artificially set "up" → probe 2 fails →
logged as up->down. Neither controller has ever reached the agent since it restarted. This is
not a degradation of a working channel; it is a wrong address read at startup. Do not read the
up-> in that log line as evidence the channel once worked.
2.2 Container → island: route-simulated, not packet-proven
The controller runs NetworkMode=bridge (172.17.0.2 on docker0, 172.18.0.4 on
traefik-public), so the guest-shell result does not automatically transfer. The image ships no
nc, wget or ip, and creating a probe container would have been a mutation, so I asked the
kernel for its forwarding decision instead — a pure lookup:
ip route get 169.254.253.1 from 172.17.0.2 iif docker0
-> 169.254.253.1 from 172.17.0.2 dev eth1 cache iif docker0
ip route get 169.254.253.1 from 172.18.0.4 iif br-<traefik>
-> 169.254.253.1 from 172.18.0.4 dev eth1 cache iif br-<traefik>
Identical on both boxes. Supporting conditions all present: net.ipv4.ip_forward=1; rp_filter=2
(loose, will not drop the asymmetric-looking reply); -A POSTROUTING -s 172.17.0.0/16 ! -o docker0 -j MASQUERADE (and the same for 172.18.0.0/16), so the source becomes 169.254.253.2;
DOCKER-FORWARD ACCEPTs -i docker0 and -i br-<traefik>.
This is strong but not conclusive — it is a routing-table verdict, not a delivered packet. It
matters because the R-50 spike chose link-local precisely because RFC-3927 "structurally does not
route" (SPIKE-island-bridge-2026-07-25.md:155-161), and that spike's P5 proved the island reachable
from the guest, never from a bridged container. The fix in §6.1 should be validated by watching
for agent channel recovered, not assumed.
3. D2 — when it started (the H3 killer)
agent_channel_unreachable, from the hub's persisted events table, which survives the
container-log truncation the brief warned about:
2026-07-25 12:31:55 drill-r50 agent_channel_unreachable
2026-07-25 12:44:19 demo-hp agent_channel_unreachable <-- FIRST
2026-07-25 12:46:42 demo-felhom agent_channel_unreachable <-- FIRST
2026-07-26 06:16:07 demo-felhom (re-fire after v0.172.0 restart)
2026-07-26 06:19:24 demo-hp (re-fire)
2026-07-26 06:25:16 demo-felhom (re-fire)
2026-07-26 06:25:41 demo-hp (re-fire)
Alongside, on demo-hp: controller_started (0.167.1) at 2026-07-25 12:42:24 — i.e. the migration
restarted the controller, which re-read the stale controller.yaml, two minutes before the first
alarm. bootstrap.json mtime is 12:42:17 on that box: the migration wrote bootstrap.json and
restarted the controller within seven seconds, and the controller ignored the new file.
VERDICT: the banner PRE-DATES today's v0.172.0 deploy by ~17.5 hours. H3 is dead.
Today's 06:16–06:25 events are re-fires: channel state is in-memory (checker.go: "No persistence —
the state is re-derived each run"), so every controller restart re-alarms. Four controller restarts
today (two per box, from the two rebuilds) explain them exactly.
4. Hypothesis verdicts
| Verdict | The observation that decided it | |
|---|---|---|
| H1 — island fallout | REFUTED | Every island leg is healthy on both hosts: felhom-agent active, bound 169.254.253.1:8443; vmbr9 up at 169.254.253.1/30 with a persistent /etc/network/interfaces stanza; guest net1/eth1 up at 169.254.253.2/30; vmbr0 untouched. TCP from the guest to 169.254.253.1:8443 is OPEN on both boxes. The island works; nothing is using it. |
| H2 — endpoint in two places | SUPPORTED — this is the cause | bootstrap.json = 169.254.253.1:8443 (mtime = migration). controller.yaml = the old LAN address (mtime predates the migration by 2–4 days). The controller dials exactly what controller.yaml says. Mechanism named in §4.1. |
| H3 — R-75 / v0.172.0 | REFUTED | First alarm 2026-07-25 12:44/12:46; first v0.172.0 deploy 2026-07-26 ~06:14. R-75 touched no agent, network or local_api surface. It is a coincidence of noticing, not of cause. |
| H4 — genuinely down | REFUTED | systemctl is-active felhom-agent → active on both hosts, since the migration, listening on the island address. The agent is up; the controller is knocking on the wrong door. |
4.1 The exact mechanism (source-confirmed)
internal/bootstrap/bootstrap.go:254:
func ensureLocalAPI(configPath string, cfg *config.Config, logger *log.Logger) *config.Config {
if cfg == nil || cfg.LocalAPI.Endpoint != "" {
return cfg // already has the agent path → nothing to do
}
MaybeIngest short-circuits for an already-configured guest (cfg.Customer.ID != "") and delegates
to ensureLocalAPI, whose contract is documented as "merges it in from bootstrap.json if
missing". It fills an empty endpoint; it never reconciles a stale one.
So: controller.yaml is generated from bootstrap.json exactly once, at Day-0, and maintained
independently thereafter. Anything that later rewrites bootstrap.json — the island migration —
silently fails to reach the controller. The mtimes are the proof: bootstrap.json 2026-07-25,
controller.yaml 2026-07-21 / 2026-07-23.
The stale comment the brief flagged is real and is a contributing signpost —
config/config.go:75 still documents the field as "host bridge IP:port, e.g. 192.168.0.162:8443",
i.e. the pre-island shape, and that is literally the value now sitting in the file.
5. Symptom B — Protected container not running: felhom-samba
Mechanism CONFIRMED.
docker ps -a --filter name=felhom-sambaon demo-hp → empty; the container was never created.- Settings:
{'enabled': True, 'server_name': 'FELHOM-HP'}— nouser_setkey, i.e. the household password has never been set. (Read only; nothing changed.) reconcileSambaAtcorrectly refuses to deploy:if !smb.UserSet { … "deploy skipped — household SMB password not set yet" … return nil }.EffectiveProtected(healthcheck.go:267) nevertheless adds it:
for _, name := range cfg.Stacks.Protected {
if name == "cloudflared" && cfg.Infrastructure.CFTunnelToken == "" { continue } // <-- has a predicate
out = append(out, name)
}
if smb.Enabled { // <-- no password / deployed predicate
out = append(out, infra.SambaContainerName)
}
Live confirmation:
[DEBUG] [monitor] Checking 5 protected containers: [traefik cloudflared felhom-controller filebrowser felhom-samba]
[DEBUG] [monitor] Protected containers missing: [felhom-samba]
[DEBUG] [monitor] Final status: fail (issues=1, warnings=0, info=5)
The diagnosis holds: sharing enabled + no password → samba deliberately not deployed → the health monitor reports FAIL for an intentional state.
A sharper finding than the brief anticipated: the function's own doc comment asserts the invariant
that is broken. healthcheck.go:260-261 says "The bring-up applies the same conditions
(stacks.EnsureBaseStack for cloudflared, ensureSamba's if !smb.Enabled { return } for samba), so
detection and deployment agree in both directions." That was true when written. reconcileSambaAt
now has a second early return on !smb.UserSet that the protected set never mirrored, so the two
sides silently diverged and the comment documents a guarantee the code no longer provides.
I caused this state. Sharing was enabled on demo-hp at ~06:20 UTC today by leg 4 of the R-75 task
(explicitly authorised by that brief). The hub data shows the box was ok for 124 consecutive
reports and flipped at the first cycle after that change.
6. D6 — blast radius
Symptom A — 4 operator emails, no customer email
notification_log, event_type LIKE 'agent_channel%':
| when (UTC) | box | channel | status |
|---|---|---|---|
| 2026-07-25 12:44:20 | demo-hp | operator | sent |
| 2026-07-25 12:46:43 | demo-felhom | operator | sent |
| 2026-07-26 06:16:08 | demo-felhom | operator | sent |
| 2026-07-26 06:19:24 | demo-hp | operator | sent |
Four, not one per cycle — the checker's alerted flag suppresses repeats within a down-spell, and
the hub applies its own cooldown (7 hub events today produced 2 emails). warning severity is
operator-only, so no customer was emailed. drill-r50 also alarmed twice on 2026-07-25, which is
consistent: it is on the same island and was the migration's first subject.
Symptom B — zero emails so far, but a customer-facing one is plausibly pending
- Zero notifications mentioning
sambaorProtected, ever. - demo-hp report health:
ok×124 through 2026-07-26 06:23:47, thenfailat 06:38:42 — one fail cycle at the time of the snapshot. demo-felhom stayedokthroughout (control: sharing there has a password and samba is running). - But the path exists and has fired before.
health_degraded/health_criticalare real notification types, andhealth_criticalhas gone to the customer channel three times historically (2026-02-24, 2026-06-15, 2026-07-18: "Rendszer állapot kritikus (volt: ok)"). A sustainedfailon demo-hp is therefore likely to produce an operator and possibly a customer email on the next transition evaluation. It had not fired as of the last snapshot. This is a live, pending consequence, not a closed one.
The English-string defect is WHOLE-SURFACE, not one-off
Every report.Issues and report.Warnings string in the codebase is English:
healthcheck.go:56 "SSD disk usage critical: %.0f%%"
healthcheck.go:82 "HDD disk usage critical: %.0f%%"
healthcheck.go:150 "Docker: %v"
healthcheck.go:169 "Protected container not running: %s"
healthcheck.go:64/87/97/115/133 "SSD/HDD disk usage high", "Memory usage high", "CPU usage high", "Temperature high"
So this is not a slip in one line — the whole health-issue surface is operator-English and is rendered on a customer-facing Hungarian page. Nothing was changed.
7. Not collected, and what it would take
| Item | Why | What would get it |
|---|---|---|
Live TCP connect from inside the controller container to 169.254.253.1:8443 |
the image ships no nc/wget/ip, and creating a probe container would be a mutation this task forbids |
one docker run --rm --network container:felhom-controller alpine nc -z 169.254.253.1 8443 — a throwaway container, no touch to the controller. Answered indirectly by the route simulation in §2.2, which is strong but not a delivered packet |
agent.json listen_addr value |
/etc/felhom-agent/agent.json read as "unreadable" on both hosts even as root — the path is likely different in v0.96.0 |
systemctl cat felhom-agent for the actual config path, or the agent's own startup log. ss already proves the effective bind (169.254.253.1:8443), so this is corroboration, not a gap in the conclusion |
Whether a health_critical email will fire for demo-hp |
only one fail cycle had elapsed at snapshot time |
re-check notification_log in an hour, or read the hub's health-transition checker thresholds |
8. Proposed fixes — DESCRIBED, NOT APPLIED
8.1 The endpoint (Symptom A) — the actual outage
Immediate, per box: set local_api.endpoint in controller.yaml to 169.254.253.1:8443 on both
guests and restart the controller. Two boxes, one line each. Safe-obvious in shape, but validate
rather than assume — §2.2's link-local routability from a bridged container is route-simulated, not
packet-proven. Success signal: [INFO] [channel] agent channel recovered plus an
agent_channel_recovered hub event.
The real fix — ensureLocalAPI (internal/bootstrap/bootstrap.go:254) reconciles rather than only
fills. Its guard if cfg.LocalAPI.Endpoint != "" { return cfg } is what made the migration a no-op
for the controller. If bootstrap.json carries a local_api block that differs from
controller.yaml, bootstrap.json should win (it is the per-guest source of truth the agent side
writes) and the file should be rewritten. Needs a small spike, not a one-liner: it changes which
file is authoritative on every boot, so it must not clobber a legitimately setup-wizard-configured or
hand-repaired guest, and the token/fingerprint must be reconciled atomically with the endpoint or a
half-updated triple is worse than a stale one. Also worth deciding whether a mismatch should be a
loud operator event in its own right — right now the two files can disagree indefinitely in silence.
Adjacent, trivial: config/config.go:75's comment still documents the pre-island shape
("host bridge IP:port, e.g. 192.168.0.162:8443"). Safe-obvious doc fix.
Process, not code: RUNBOOK-island-migration.md should gain a post-migration verification step
that asserts controller.yaml and bootstrap.json agree, and that the channel recovered. The
migration reported success while leaving every controller on the fleet blind to its agent for 18 h.
8.2 EffectiveProtected (Symptom B)
internal/monitor/healthcheck.go:267 — gate samba on the same predicate the bring-up uses, i.e.
smb.Enabled && smb.UserSet, mirroring the cloudflared branch five lines above. Safe-obvious, one
condition, and it makes the function's own doc comment true again — that comment at L260-261 must be
updated in the same change, since it currently asserts a parity that does not hold. A red-proof is
easy: enabled-without-password must yield a protected set excluding samba, and enabled-with-password
must include it.
8.3 The English strings on a Hungarian page
Whole-surface (§6), so this is a copy/i18n arc, not a one-line fix — every report.Issues and
report.Warnings producer, plus a decision about where translation belongs (at the producer, or at
the render/notification boundary where operator-English and customer-Hungarian already diverge). The
health_critical customer email suggests the boundary already exists and these strings simply bypass
it. Needs a spike to pick the seam; the strings themselves are then mechanical.
9. Observations — noticed, not acted on
drill-r50is also alarming (2026-07-25 10:04, 10:11, 10:14, 12:31) and was the migration's first subject. If that VM still exists it is likely in the same stale-endpoint state. Out of scope here; worth a look when the fix lands. Its 10:04–10:14 alarms precede the demo boxes' 12:4x, so it may also carry earlier, unrelated spike churn.expected_backup_missedis firing nightly on all three customers (2026-07-26 03:00, including a customer-channel email to demo-felhom: "newest backup is 176h0m0s old" = 7.3 days). That is a separate, louder problem than either symptom here, and it reached a customer channel. Almost certainly downstream of the same outage — the agent channel is how PBS backup state is learned — but I did not verify the causal link.- The channel checker has no persistence by design, so a restart re-alarms and the operator sees a "new" alert for an old condition. That is defensible (state re-derivation is the model) but it is why this looked like it started today.
rp_filter=2(loose) on both guests is what makes the container→island path viable at all; under strictrp_filter=1the asymmetry could bite. Worth knowing before anyone hardens it.- demo-hp has been up 2 d 12 h and demo-felhom 3 d 13 h — neither has rebooted since the
migration. So the persistent
vmbr9stanza in/etc/network/interfacesis present but has never been exercised across a boot on either box. The island's reboot-survival remains untested in the field. - Sharing is still ON with no password on demo-hp, which is what keeps Symptom B live. Left
exactly as found, per the prime directive — but it means the box will keep reporting
failuntil either the fix lands or sharing is turned off.