CHANGELOG + REPORT + CONTEXT: v0.133.0 released (tag + package verified by download), not delivered
gates / gates (push) Successful in 13s
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,65 +1,29 @@
|
||||
# REPORT — agent v0.132.0: a controller that dies slowly is reported, not just restarted
|
||||
# REPORT — agent v0.133.0: a restore-test can never fill a box's disk (2026-09-24, R-672, R-673)
|
||||
|
||||
**2026-09-17.** R-539, operator ruling 3 of 2026-09-16. Architecture: `felhom.eu/documentation/architecture/03-host-agent.md` (the supervisor and its budget).
|
||||
Full record: `felhom.eu/documentation/audits/r672-2026-09-24/README.md`.
|
||||
|
||||
## Claims in the brief that turned out wrong — named first
|
||||
## Shipped (released, not delivered)
|
||||
Tag `v0.133.0` (`9bdb4da`), package sha256 `3aa30345…e69b6`, verified by anonymous download. Delivery needs the
|
||||
operator's signed `agent_update` job per box.
|
||||
|
||||
1. **„Emit `controller_slow_crashloop` from the agent."** The agent has no event channel. It carries
|
||||
timestamps in its heartbeat stanza and the **hub** mints the event when one moves. So the build was
|
||||
three wire fields here **plus** a hub checker change (hub v0.117.0), not an event registration alone.
|
||||
2. **„Five kills 20 min apart is too long — make the interval configurable for the test."** Not needed,
|
||||
and not done: the proof restart from B.4(a) was already persisted, so four more real kills ~8 minutes
|
||||
apart reached five in 24 hours inside the session, with the **production** window. 8 minutes keeps any
|
||||
15-minute window at two restarts, so the fast brake never interferes.
|
||||
- **Space preflight** before anything is created: free ≥ restored × 1.2 + 5 GiB, `restored` UNCOMPRESSED (vzdump
|
||||
log "Total bytes written" / PBS snapshot size), thin metadata with room, off the tested guest's pool when
|
||||
another eligible storage fits, unknown refuses, reported as a non-pass (`skipped`).
|
||||
- **Scratch teardown retried every 10 min**; operator told once after 3 failed tries.
|
||||
- **Thin pool ≥ 90 %** → immediate host report (hub v0.124.0 alarms `storage_fill_critical`, per pool per 6 h).
|
||||
- **R-673:** the stale-lock sweep on the same timer, under the one-heavy-op gate.
|
||||
|
||||
## What shipped
|
||||
## Red-proofs (each seen failing; files in `felhom.eu/documentation/audits/r672-2026-09-24/redproofs/`)
|
||||
1. preflight removed → the 2026-09-24 restore issued again; 2. archive FILE size used → "the compressed file size
|
||||
was used"; 3. timer pass a no-op → "the leaked scratch was not destroyed by the timer"; 4. sweep without the gate →
|
||||
"the sweep ran while a backup held the gate"; 5. no 90 % edge → "0 report requests, want 1"; 6. every skip dropped →
|
||||
"a space refusal never reached the host report". Full suite `go test ./...` green; `agent_gates.py` OK.
|
||||
|
||||
`internal/localapi/controllersupervisor.go`: beside the unchanged 3-in-15 brake, restarts the supervisor
|
||||
performed in the last 24 hours. At the fifth, `slow_crashloop_since` moves (at most once per 24 h);
|
||||
`slow_crashloop` and `restarts_24h` ride the stanza (`internal/hub/report.go`). Persisted per guest in
|
||||
`/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json` (tmp + rename, 0600); unreadable or
|
||||
corrupt → WARN and a clean start. Never stops restarting. Deliberate kills count. The startup line prints
|
||||
the new limits.
|
||||
|
||||
## Red-proofs — each seen failing, then passing
|
||||
|
||||
| test | break | failure seen |
|
||||
|---|---|---|
|
||||
| `TestControllerSupervisor_SlowCrashloop` | no counter | `five restarts 20 minutes apart did not raise slow_crashloop — this is R-539` |
|
||||
| same | once-per-24h guard removed | `the raise moved again on the 6th restart … mailed per restart` |
|
||||
| `…SlowCounterSurvivesAgentRestart` | save removed | `the agent restart reset the slow counter … Restarts24h:1` |
|
||||
|
||||
Negative control: restarts 7 h apart never raise it. Wire shape extended (`restarts_24h`, `slow_crashloop`).
|
||||
|
||||
## Gates and release
|
||||
|
||||
`go build ./... && go vet ./... && go test ./...` — green, 30 packages. `agent_gates.py --fast` — all OK.
|
||||
Released by `scripts/release-agent.sh`: tag `v0.132.0`, sha256
|
||||
`4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321`, verified by independent download.
|
||||
**Not vouched** for Day-0 installs (the operator's act). Code `18d03bd`, release record `CHANGELOG` commit after it.
|
||||
|
||||
## Delivery — operator-signed, per box (ruling 1 of 2026-09-16)
|
||||
|
||||
| box | signed | authorised → completed | committed | startup line |
|
||||
|---|---|---|---|---|
|
||||
| demo-hp (`demo-hp-bb76ea`) | 08:27:36Z | 08:29:25Z | 08:30:29Z | `slow_crashloop_max=5 slow_crashloop_window=24h0m0s` |
|
||||
| N100 (`demo-felhom-8363b5`) | 08:27:36Z | 08:34:48Z | 08:35:52Z | same |
|
||||
|
||||
Peti's box: not touched. Evidence: `felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partC-agent-update.txt`.
|
||||
|
||||
## Live validation — PASS, production window
|
||||
|
||||
On demo-hp guest 9201: five controller kills (08:51:53Z from B.4(a), then 08:57, 09:05, 09:13, 09:21),
|
||||
each restarted by the agent in 42–56 s, never by hand; the fast brake never armed. At #5:
|
||||
`SLOW CRASH-LOOP — … restarts_24h=5 window=24h0m0s threshold=5`; persisted file holds the five times and the
|
||||
raise. The hub minted `controller_slow_crashloop` at 09:29:40Z and **exactly one** operator mail arrived.
|
||||
Evidence: `…/partC-live-slow-crashloop.txt`.
|
||||
|
||||
## Teardown
|
||||
|
||||
Nothing provisioned. Stated effect on the demo box: 9201's counter holds five restarts and stays raised
|
||||
until 09:22Z on 2026-09-18; it cannot mail again inside that window.
|
||||
## Live (demo-hp, the on-demand `--selftest=restore-test`, cadence OFF)
|
||||
(a) factor 10 → refused: needs 215.5 GiB, has 22.1 GiB. (b) normal margin → refused: restoring 21.1 GiB needs 30.3 GiB,
|
||||
has 22.1 GiB — correct: no full restore-test fits demo-hp under 80 % pool use. (c) a forced teardown failure could not
|
||||
run live (it needs a scratch guest, which (b) shows cannot be created within the 80 % rule) — unit tests only. Pool
|
||||
58.99 % before and after every run; nothing created.
|
||||
|
||||
## Observations
|
||||
|
||||
1. The persisted file writes the zero raise time as `"0001-01-01T00:00:00Z"` (Go `omitempty` does not omit a zero `time.Time`). **NOT-A-FINDING: the loader reads it back as zero (`IsZero`), which the live file and the restart test both exercised; cosmetic only.**
|
||||
1. demo-hp has no second eligible storage for a restore-test: `nvme-scratch` takes `rootdir`, but the agent holds no grant there. FILED: R-672
|
||||
|
||||
Reference in New Issue
Block a user