910fd911244c113b259691ddef69090716ab501a
596 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
910fd91124 |
agent 0.130.0 published and vouched; R-347 closed, R-349 + R-350 filed
gates / gates (push) Successful in 14s
Released via scripts/release-agent.sh: tag v0.130.0 at 7569f34, sha256
a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3,
verified by independent download and reproducible byte for byte with
-trimpath -buildvcs=false.
Vouched agent 0.129.0 -> 0.130.0 in the Day-0 manifest. Only the agent
fields changed: min_agent stays 0.129.0 because it states what the GOLDEN
CONTROLLER requires, and raising it would have HELD the floor for every
box below 0.130.0. Global floor untouched at 0.216.0 -- and on hub
v0.106.0 it is a separate form with its own action, so publish-train
rule 2's hazard no longer exists in the shape its incident describes.
No --no-verify: the CHANGELOG heading was flipped only after the tag and
package existed, so release-complete passes on the real artifact.
R-349: the fleet was running a DIFFERENT binary under the same version
name -- the proof deploy was a hand build, the release is -trimpath.
Self-update could never have corrected it, because every version check
compares the string. Both boxes reinstalled from the downloaded package.
The proper fix exists in miniature as wrapper_sha256 and was never
extended to the agent's own binary.
R-350: I printed the hub password into the session transcript via
curl -w '%{redirect_url}' -- the hub answers 303 and curl re-attaches the
credential. Not in git, not in any committed file, not in the evidence
directory. Rotation is the operator's call.
ep0 closes at fd 17, ESTAB 0, CLOSE-WAIT 0 -- its t0 baseline -- and was
read-only for this entire arc.
|
||
|
|
57dd62b097 |
R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed. |
||
|
|
9299f85c4b |
SPIKE ep0 connections: CI green by run id (360/237, 19672e685)
gates / gates (push) Successful in 14s
|
||
|
|
19672e685e |
SPIKE ep0 connections: the leak is felhom-agent's, not the poll rate
gates / gates (push) Successful in 14s
R-341's first dated check, taken at +46.2 h: fd 17 -> 405 over 166,251 s
= 201.6/day. Pre-registered range was 370-450; observed 388. UNCHANGED,
as predicted. CLOSE-WAIT is 0 -- absent entirely, not merely flat.
Q1: exactly two peers, 194 each, no third party.
Q2: outcome (a). ep0 388 = 194 + 194 on the boxes, twice, and the four
new sockets carry the same source ports on both sides. 0 closed in 31 min.
The finding: all 388 are held by felhom-agent. pvestatd and
proxmox-backup-client made 162,404 requests and leaked zero. Mechanism is
a per-cycle http.Transport with a zero-value IdleConnTimeout that nothing
ever closes (internal/pbs/client.go:56, main.go:1486). R-336's premise
does not survive this -- cutting the poll rate would have fixed nothing.
Q3 NOT measured: Phase C held at STOP 1, prediction pre-registered first.
Part 0 captures the due-checks gate's first conviction on a real overdue
date (rc=1, names R-341, sole failure among 10 gates). Row cleared at
Part 4, after the result was recorded in R-341, not to make a push work.
New: R-344 (the transport leak), R-345 (hub/Makefile pushes :latest),
R-346 (ActiveEnterTimestamp reads 5h56m early -- NRestarts is still 0).
|
||
|
|
ab2262c91c |
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
|
||
|
|
0a5e9b14dc |
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10. |
||
|
|
f267bc047f |
R-334 CLOSED, quoting the green CI run id
gates / gates (push) Successful in 14s
Golden 0.216.0 baked, published and vouched; the row is closed with CI run
353 (head_sha
|
||
|
|
7d81681d6e |
golden 0.216.0: baked, published, vouched — gates green again
gates / gates (push) Successful in 13s
Closes the two-release day-0 gap that has been convicting CI since 2026-08-14. Run against RUNBOOK-manual-build.md 4.0 + 4.1. GOLDEN_VERSION = 0.216.0 GOLDEN_SHA256 = ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b archive = 656,970,239 bytes, controller image 0.216.0 template = debian-13-standard_13.6-1_amd64.tar.zst (listed live, not reused) Baselines re-read on the machine and all four matched the sheet: controller v0.216.0, its MinAgent 0.129.0, agent v0.129.0, previous golden 0.214.0. The published agent artifact for the vouched agent_version was confirmed present in the package registry rather than inferred from a CHANGELOG, and the R-216 check passed on the machine: MinAgent is EQUAL to, not above, the newest published agent. Verified beyond the script's own claim: the artifact was downloaded back out of Gitea and hashed, and it matches GOLDEN_SHA256 exactly. A script printing a digest and the registry serving those bytes are two different claims. Pass markers (corrected post-R-233 list) all present, quoted with line numbers in pass-markers.txt; excluding/FATAL absent; there is no mp1. Token never reached a command line: copied file->file, read inside the VM by the runner. systemctl show grep = 0. Token-leak grep on the COMMITTED log run with its positive control FIRST -- seeded copy 1, real log 0 -- because a grep -c that matches nothing also returns 0. Teardown: guest destroyed and purged, token/runner/script/log shredded AFTER the log was copied out, qemu exit confirmed with ps -eo comm (not pgrep -f), disk reverted to virgin. Vouched by the operator; verified by reading the hub's own store: golden 0.216.0 / agent 0.129.0 / min_agent 0.129.0, and the hub's recorded sha256 matches the independently downloaded artifact. That check was necessary because golden_currency_gate.py says of itself that it checks the BAKE, not the vouch. repo_gates.py --fast now rc=0, all nine gates OK -- first fully green run since 2026-08-14. Capability map deliberately NOT changed: the day-0 row cites drill documents, and the map's only golden literal is a dated historical citation on the recovery-journey row which bumping would falsify. R-334 is closed in a follow-up commit quoting this push's CI run id, since closing it without one would leave the ambiguity a third time. |
||
|
|
3e50902a98 |
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
|
||
|
|
435e044cf1 |
INCIDENT/R-336: the leak is measured live, not assumed
gates / gates (push) Failing after 15s
Post-restart baseline on ep0: 19 fds at 16m43s (from 18), 1 CLOSE-WAIT, Recv-Q 0. One descriptor per ~17 min is ~85/day, which agrees with the ~73/day implied independently by the failure itself (1016 sockets over 14 days of uptime). Two estimates of the same slope agreeing turns "the ceiling raise is mitigation, not a cure" from a plausible claim into a measured one, and puts the next ceiling at ~2 years instead of a fortnight. Recorded because standing rule 3 asks for a positive observable: this is it, and it fired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
ebfd0967c1 |
INCIDENT + registers: ep0's PBS proxy served nobody for 9.5h (R-336..R-338)
gates / gates (push) Failing after 12s
Two whole_guest_backup_failed alerts at 04:30 and 04:32 CEST were one incident, and not on either customer box: ep0's proxmox-backup-proxy was active, holding its listening socket, and accepting nothing. Root cause: accept() returning EMFILE. The process held exactly 1024 fds -- its systemd-default soft RLIMIT_NOFILE -- of which 1016 were sockets and 547 connections sat in CLOSE-WAIT. The 1024-deep accept backlog had overflowed (Recv-Q 1025), so every client timed out. It was wedged from its own loopback too, which is what moved this from a network problem to a process problem. Fed by ~85k requests/day (a flat 3,538/hour) against an endpoint written to weekly, that leak reached the ceiling in 14 days of uptime. Fix: LimitNOFILE=65536 drop-ins for both PBS units, restart, verified from both boxes (200 in ~0.1s, felhom-pbs active), then re-drove the missed backups through the product path -- POST /backup?target=felhom-pbs on each agent's local API, not a hand-run vzdump. demo-felhom ct/9201/2026-08-18T03:57:43Z 4.10 GB 36.4s demo-hp ct/9201/2026-08-18T03:58:43Z 4.29 GB 41.5s Both host reports now carry felhom-pbs success=true, so the hub is green on the evidence rather than on a restart having been performed. No data lost, no backup skipped: the daily local tier was never affected and the PBS tier is weekly, so the window cost exactly one attempt. Evidence copied off ep0 BEFORE the restart, per standing rule 5. Filed: R-336 (the ~1 req/s poll rate is the real defect; the raised ceiling is mitigation, not a cure), R-337 (a status endpoint that trailed its own artifact by minutes then caught up -- WATCHING, downgraded from the defect I first wrote, because it self-corrected), R-338 (demo-hp is not on the R-50 island at all and nodes.md says it is; its local API is bound to the customer LAN). R-334 updated: still open, now one version wider (controller 0.216.0 vs golden 0.214.0). golden-currency is the only failing gate and is inherited -- it reads files this session did not touch -- so this push used --no-verify, stated per .claude/rules/gates.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
ea16a21bff |
docs(register): narrow R-332 — restart persistence is now proven live
gates / gates (push) Failing after 14s
The 0.215.0->0.216.0 redeploy replaced the container and the state file came back with the previous version's changed_at, so the new container loaded the pre-restart record rather than re-baselining. The verdict path itself, and the already-alerted-disk restart case, remain unproven. |
||
|
|
fa4748d4dd |
docs(register): R-335 — one physical disk walked twice per run, sustained against itself
gates / gates (push) Failing after 16s
Found on live hardware after the v0.215.0 deploy by reading the check's evaluated count against its own persisted state file. Closed in v0.216.0. |
||
|
|
767960bb11 |
docs: disk-health phase 1 — capability map, roadmap arc, register rows R-328..R-333
gates / gates (push) Successful in 14s
- capability-map: the disk-failure scenario no longer says a failing disk has never been seen. Healthy path + delivery + the severity wire stay PROVEN-LIVE; the new Hiba-from-counters path is IMPLEMENTED and explicitly NOT proven-live (R-332), because it has only ever run against the fixture's values. - ROADMAP R-73: phase 1 shipped; the genuinely hub-side half splits into R-330 (phase 2, a declared wire change under G-1) and R-331 (phase 3, growth-rate detection and retiring the static 64). Its premise 'no demo hardware exposes real SMART' is retired — a real failing drive is now committed as a fixture. - register: R-328 (the severity drop, CLOSED and proven live side by side), R-329 (app_start_failed has the same defect, needs a decision first), R-330/R-331 (phases 2 and 3), R-332 (the Fail path has never fired on real hardware, WATCHING), R-333 (NVMe temperature bands measured 2 degrees from tripping on a healthy drive; and the agent's smartctl has no -n standby). |
||
|
|
848de8153d |
docs(audits): first genuinely failing disk — the SMART PASSED trap
gates / gates (push) Successful in 12s
Commit the raw evidence from ST3000VX010 S/N Z6A07P2G (/dev/sdg on DooPlex), which went 8 -> 352 unreadable sectors 11-13 Aug while smart_status.passed stayed true throughout. - fixtures/smart-ST3000VX010-failing-2026-08-14.json: raw smartctl -a -j, verbatim - fixtures/smartd-history-sdg-2026-08-14.txt: 406 smartd journal lines, 11-14 Aug - DIAG-smart-passed-trap-2026-08-14.md: the mechanism (attrs 187/197/198 all carry thresh 0, so a normalized value that floors at 1 can never fail the overall verdict on unreadable sectors), the non-monotonic timeline, the three controller defects with locators, and the counterfactual: zero emails would have been sent. |
||
|
|
e0b56c976f |
REPORT + CONTEXT: the third name, the second door, and a number that answered a different question
gates / gates (push) Successful in 15s
Three rules carried forward. A name must separate on the STEM, not the noun — naming this secret after the act it is used in would have recreated the trap, because the other factor on the same page is the „Párosító kód". A guard is worth what its positive control is worth: this one's selftest convicted its own step-3 case and found a defect in the guard itself. And a suppression must rest on the machine's own declaration, then be checked for the SECOND door — recording the disabled state rather than deleting it is what let the deadline check skip it too. Yesterday's report is preserved to audits/ because it carries the only record of the self-heal verdict (Part C was dropped, so that reasoning is in no register row) — the rule written last night, applied to itself the first time it mattered. |
||
|
|
b03a105375 |
hub v0.105.0: the third name, a machine told to be quiet, and a guard for the hub's own words
gates / gates (push) Successful in 17s
Hub only. No controller change, no agent change, no wire change — nothing to bake. demo-hp untouched: the operator is re-deploying it this evening. R-323 — the five-word phrase is „Tulajdonosi jelmondat". It was „Visszaállító jelszó": one word from the name retired last week, and false besides — it restores nothing, it proves the account owns the box being bound. Five sites, all in the hub; felhom-controller and felhom-agent carry the name nowhere, so no halt and no bake. Both suggested names were rejected with reasons: „Fiókjelszó" would collide with the dashboard login (a DIFFERENT real secret), and „Összekötési jelszó" would leave the two factors on this page separated only by kód-versus-jelszó — the exact shape being removed, since the other factor is the „Párosító kód". The chosen name differs on both axes, stem and noun. Naming only; the acceptance pin drives the real handler. R-324 — the hub's customer copy is under a guard for the first time. Retired names banned across all 95 hub files; retrieval stems registered in four declared customer surfaces. The selftest found a defect in its own instrument on the first run. One shared vocabulary in scripts/, drift-checked into the controller gate rather than copied (R-325 removes the scaffold). R-321 — a machine we told to be quiet is no longer reported as dead, and it was two doors, not one: because the state is RECORDED rather than deleted, the morning deadline check can skip it too. A deleted state returns "", which is not "down" — R-195's shape returning through a second door. The clock runs from the report the hub can see, so re-enabling starts it there and emits no recovery for an outage that never happened. Three red-proofs; the one that matters showed a genuinely dead machine sitting at "disabled" when the suppression was made unconditional. R-326 — "which claims are unproven" is answerable by a command now. The nine I have been repeating was the count of claims the 9 August pass DOWNGRADED, not the count of unproven ones. The real figures: 55 claims, 23 walked, 32 not — and only 6 of those 32 cite evidence. Its first run found a stale claim (R-327). |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
2d05b29b82 |
REPORT: installer v1.28.0 published and verified live; evidence, and the Part 1 logs I lost again
gates / gates (push) Successful in 13s
|
||
|
|
fc737b0fc0 |
installer v1.28.0: the removal genuinely reverses the installation (R-316)
gates / gates (push) Successful in 13s
v1.27.0's fix worked exactly once per machine. Measured on drill-r50 from virgin, on the PUBLISHED v1.27.0, before anything was changed: cycle 1 recorded 'no' and freed :53; cycle 2 recorded 'yes' and left dnsmasq running on 0.0.0.0:53; cycle 3 refused, exit 1. Every box already in the field is at cycle 2, and a reinstall onto a machine that has had Felhom is cycle 2 by definition. Why cycle 2 says yes: the preflight's ownership question is dpkg-query package presence and nothing else - not the absence of a record. Stopping the unit and leaving the package made our own package read as the household's one cycle later. Now the uninstall removes the package when the record says we installed it. Order unchanged and load-bearing: read the record, act, then delete the state file that holds it. TWO packages are recorded, because dnsmasq ships the unit and dnsmasq-base ships /usr/sbin/dnsmasq, and each is taken back only if we added it. The dependency check is a SIMULATION, not a guess: apt-get -s purge is asked what it would remove and the purge proceeds only if that set is a subset of ours; otherwise stop+disable, naming the package that blocked it. Never interactive, never fatal, and the success is re-queried rather than read off an exit code. Watched: three fixed cycles -> install 3 PASSES; a household resolver untouched; a dependent package not purged and named; no record -> untouched with the command named. Red-proofs with the mutation asserted applied: remove the purge -> cycle 3 refuses in those exact words; remove the ownership check -> a household resolver is purged; infer ownership -> the guess is taken. Also: R-317 (the agent stats a path dnsmasq-base owns to decide whether to install dnsmasq - pre-existing, now reachable), R-318 (no honest ownership marker exists for existing boxes; the preflight message is the mechanism), and the status page's decisions section rewritten to say what each decision costs and what doing nothing selects. |
||
|
|
d102ca5767 |
Vouch delivered fleet-wide; the MinAgent hold observed firing and releasing for the first time
gates / gates (push) Successful in 20s
Operator vouched golden 0.214.0 / agent 0.129.0 / min agent 0.129.0 and raised the floor to 0.214.0. Both artifact shas in hub_settings match the bake and the release byte-for-byte. With the floor at 0.214.0 and demo-hp still on agent 0.128.0, the hub HELD the controller floor - 'agent 0.128.0 < MinAgent 0.129.0 (controller floor withheld)' - and released it 6 seconds after the agent was brought up. That is the R-216 guard doing exactly what it exists for, seen firing for the first time, and it is the argument for declaring MinAgent in the CHANGELOG header. Both demo boxes verified on the boxes: controller 0.214.0, agent 0.129.0, healthy. |
||
|
|
684cd2eb11 |
R-265 third sighting: the jobs API names the failing step when the log 404s, and a re-run disambiguates
gates / gates (push) Successful in 15s
|
||
|
|
4906aeb3f9 |
R-311 proven live (HTTP 422 on hardware); R-308 WITHDRAWN — my quoting bug, not a stale credential
gates / gates (push) Successful in 19s
The live test read as a FAILURE for twenty minutes because I stripped only double quotes from a credentials value wrapped in SINGLE ones, sending a literal ' as part of the recovery code. Correctly unquoted: the old code returns 422 with opens_retained=true and the supersession date; a wrong code still returns 400. The same bug produced the R-308 finding in the previous report. The dashboard password is fine - HTTP 302 with a session cookie on the first try. Third time this project has produced a wrong 'the credential is stale' verdict from that one trap. |
||
|
|
8b188bea68 |
hub v0.103.0 — a host can read the packages we kept for it (R-311)
gates / gates (push) Successful in 37s
ListSupersededEscrow had zero production callers for nineteen days. It is the only reader of a retained identity_blob, so the retention shipped in v0.93.0 was material the product could not reach - proven on the fixture 2026-08-12, where a code that opens a retained package was answered as a code that opened nothing. New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row GET, same recovery-mode gate, same audit event written BEFORE the bytes leave, capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count - they retain the PBS key, not the repository password, so they can never open what the caller is asking about, and serving them would let the screen claim an earlier package is openable on exactly the boxes the original defect hurt. The count is returned because their existence is load-bearing and underivable by the caller. The trade, stated rather than waved through: the hub still cannot read any of it - sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held. What widens is volume, bounded by self-scope, the recovery-mode gate and the cap. The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it; the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than 174. A positive control shows that check is name-presence, not decodability - filed as R-315 rather than reported as coverage. Also: golden 0.214.0 baked, published and round-trip verified; the countdown on demo-felhom cancelled on the operator's ruling (R-307); the spike that halted Part 3 recorded as R-312; the set-aside store found unrecoverable as R-313. Six hub tests through the real endpoint; four red-proofs asserted applied. |
||
|
|
6362bb6cb6 |
hub v0.103.0 — a host can read the packages we kept for it (R-311)
ListSupersededEscrow had zero production callers for nineteen days. It is the only reader of a retained identity_blob, so the retention shipped in v0.93.0 was material the product could not reach - proven on the fixture 2026-08-12, where a code that opens a retained package was answered as a code that opened nothing. New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row GET, same recovery-mode gate, same audit event written BEFORE the bytes leave, capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count. They retain the PBS key, not the repository password, so they can never open what the caller is asking about; serving them would have the agent try packages that cannot succeed and would let the screen claim an earlier package is openable on exactly the boxes the original defect hurt. The count is returned because their existence is load-bearing and underivable. The trade, stated rather than waved through: the hub still cannot read any of it - sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held. What widens is volume, bounded by self-scope, the recovery-mode gate and the cap. The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it; the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than 174. A positive control shows that check is name-presence, not decodability - filed as R-315 rather than reported as coverage. Six tests through the real endpoint; four red-proofs asserted applied. |
||
|
|
c1319a91a8 |
Correct the drill's wall clock to the measured end time (17:45, not the estimated 17:55)
gates / gates (push) Successful in 19s
|
||
|
|
1d5f2b8bb6 |
DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption survived a week. (a) The material IS retained. host_escrow_superseded id 11 is the first retained row in fleet history to carry identity_blob (572 B), byte-identical to the pre-supersession row (sha256 a10032341c8584ed...). (b) The retained material DOES open the old store. Unsealed with the old recovery code it yielded a password byte-identical to the pre-change one, and restored three planted files byte-identical from a store the box itself could no longer open - including a Hungarian accented filename verified as raw bytes. Negative control ran first and failed closed. (c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero production callers; the recovery path selects FROM host_escrow. Asked with the code that had just worked by hand, the product answered "the recovery code did not open the sealed bundle". A valid code for retained history is reported as a bad code - the R-224 class again. R-304, rank 1. Both installer faults were watched happening first, so installer-v1.27.0 is now published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller 0.98.3 against a vouched 0.213.0, below the floor and below the version carrying the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own next install refused. R-297 and R-300 CLOSED. Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns on the second reinstall, proven), R-306 (--preflight-only writes state it says it does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 - operator decision), R-308 (stored controller password stale), R-309 (the day-0 runbook's publication claim has been false since R-110), R-310 (two edges). Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked operator-only. Phase A logs did not survive the intermediate revert; recorded. |
||
|
|
fbe1155fbb |
R-302 docs: register rows, the two rules earned twice, STATUS
gates / gates (push) Successful in 27s
Closes R-296 (verified: shipped in v0.212.0) and R-301 (premise confirmed, fixed in v0.213.0). Files R-302 with WHY the obvious condition was rejected, and R-303 for the missing markOrphaned guard - the co-render is now harmless, not impossible. Bake evidence for golden 0.213.0. |
||
|
|
890a474ff2 |
STATUS back to one screen; close R-280 and R-294
gates / gates (push) Successful in 21s
211 lines -> one screen. Moves closed items out, corrects the tester paragraph, states the floor situation as the operator's one-field call, and stops asking him to decide something that shipped. |
||
|
|
125aec1be2 |
R-300: uninstall no longer leaves dnsmasq blocking the next install
gates / gates (push) Successful in 18s
Removing the snippet and restarting left dnsmasq enabled and unconstrained on 0.0.0.0:53, so the next byo install's preflight refused and the customer went debugging a home network that was never at fault. Ownership is recorded at preflight (the only moment it is a fact - the package is installed by the agent, not this script) and honoured at removal. Boxes already in the field carry no record and fail safe to restart-only, with the reason and the command logged; the preflight message covers them instead. Not observed live - no installer-v1.27.0 tag is cut. Files R-299..R-301. |
||
|
|
238954405c |
golden 0.212.0 baked and published; bake evidence
gates / gates (push) Successful in 23s
GOLDEN_SHA256=4b0a7dacc503c38732ed0a44949398639248c7fbd90758a1e4a047c21a7a15d8 Round-trip verified on the served bytes. Not vouched - the operator's. |
||
|
|
0bdffbe865 |
SPEC correction: surface 1 was NOT accurate, and the guard matched one inflection
The spec listed backups_remote.html:98 as 'Accurate; keep'. It ended with the same unevaluable promise as surface 2, in the plural - and because the spec's own guard was written against the singular form, it could not catch it either. Both corrected; implemented in controller v0.212.0 (R-299). |
||
|
|
f76cbf0ec7 |
Part 0: correct the PETI record - the mitigation it named does not exist
gates / gates (push) Successful in 20s
The row said a drive failure there means offsite-only recovery. Re-read from the hub's own store: no host row (deleted 2026-07-15 08:56:22, escrow_acked=0), no escrow of any kind, and offsite backup never ran once (escrow_state pending, snapshot_count 0 - the fork-4 guard working, not a fault). The local app-data repo was empty too and the whole-guest vzdump shares the failing device. If that drive fails today, everything on it is lost. Records the fact and leaves the parked/not-parked ruling open - that is the operator's call and does not need restating to be true. STATUS.md no longer reads the absence of a hub record as reassuring. |
||
|
|
999b0a35f8 |
golden 0.211.0 baked and published; bake evidence
gates / gates (push) Successful in 22s
GOLDEN_SHA256=8593516889eb93fe1691410d7306be8cb87ee835b8d2378740eb34022272f849 Round-trip verified on the served bytes. Not vouched - that is the operator's. |
||
|
|
eb600872f2 |
R-297: installer compares a local golden against the manifest before using it
Step 7 short-circuited on any local golden archive with no version compare, no digest and no warning, so the manifest sha256 was consulted only on the fetch path. Local discovery is newest-by-filename: correct by recency, never by verification. A box could reinstall from a stale archive and come back below the version where the offsite recovery screen exists. Digest first, then the baked controller tag. An auto-discovered mismatch re-fetches the vouched golden; an operator-named mismatch refuses. An unreadable manifest refuses rather than passing. Not published: installer-v1.26.0 is deliberately not cut until a fresh install has been observed taking a stale local golden on drill-r50. Also files R-295..R-298. |
||
|
|
11a5c3bd92 |
R-265 second sighting: a CI run failed and its log cannot be retrieved
gates / gates (push) Successful in 13s
felhom.eu run 293 ( |
||
|
|
c04f933d0b |
The census answers no, three receipts found, and the prune was on file all along
gates / gates (push) Successful in 13s
CENSUS (read-only, hub store, tester's machine not contacted): no machine that is not ours can be in the state that cost demo-felhom its history. The hub holds escrow for three hosts; both demo boxes lost their pre-fix key in the same four hours on 2026-08-04; peti-felhom and david have no host row and no escrow at all. A control ran FIRST and had to pass -- the query returned "present (572 bytes)" for a host known to have material and "absent (NULL)" for one known not to. Corrected my own instrument on the way: a date-only comparison mislabelled both losses as after the fix, so the in-force moment is now pinned from the hub's first post-fix escrow row (11:11:37Z), which independently agrees with the register. PART 1 ESTABLISHED. The prune is recorded inside R-267 -- the row about the Configuration page being slow -- because pruning artifacts is what made that page fast. Arithmetic checks (23+7=30, plus three versions that only surfaced after the first thirty moved them onto page one = 33) and the PAGINATED listing shows both generics at exactly ten. R-291's blocking condition is released: the operator was being asked to establish something already written down. And my counter-argument yesterday was wrong in exactly the way R-267 warns about -- "containers hold 19" came from an unpaginated query; paginated they hold 270 and 169. RECEIPTS: three restored (drives.enrol, backup.tier1, fail.lost-recovery-code), each citing the document that walked it; the map already read PROVEN-LIVE for all three, so this follows the map rather than raising a status in the view. NINE HONEST GREYS. fault.selfheal's best hit argues against it -- an incident recording self-heal's absence through a 1h15m outage. THE DECAY RULE FIRED FOR THE FIRST TIME. backup.restore-proof has a receipt from 28 July and is superseded anyway: demo-hp's restore-test failed 5 August and the box has since been rebuilt. A claim about a continuing behaviour cannot rest on an old observation. The capability map still reads PROVEN-LIVE and is now the thing out of step -- recorded, not silently rewritten. PART 4 specified, not implemented. The orphan card promises restorability the box rendering it cannot evaluate: the discriminator is on the hub and no wire field carries it. A conditional promise the system cannot evaluate is the same defect as an unconditional false one, so the copy stops promising, says what happens, and names a route. Ships with the next controller change so one bake covers both. |
||
|
|
67eced8fbf |
demo-felhom is protected again, and the authorised recovery could never have worked
gates / gates (push) Failing after 10m40s
Checked before acting, and the check is the finding. The box's local key and the hub's sealed escrow key hash to the SAME value (c60c8bc737a6b7c6...), and that key answers "wrong password or no key found" against its own repository. Running the recovery would have returned a key the box already held and which was already proven not to open the store. The store was written under 48741892f0ef4d59... -- host_escrow_superseded id=4, superseded 2026-08-04 07:20:08, identity_blob NULL. The restic password lives only in the identity bundle (escrow/identity.go:39, read by recover.go:91), so it is unrecoverable by construction; the surviving K-escrow payload is 64 bytes, a wrapped key, far too small to carry it. Same shape the register already records for demo-hp, four hours the wrong side of the retention fix. Took the operator's stated fallback instead: the orphan reset through the customer's own card. Old store moved aside, never deleted, to /home/felhom-repo.orphaned-20260810 (1.2 GB); fresh repository under the current key; offbox_repo_reset audited hub-side. Then PROVEN rather than assumed -- last_status ok, 10s, and the snapshot's CONTENTS listed: opengist compose files, manifest.json and volume-dumps/opengist_opengist_data.tar. Not an empty backup calling itself successful. R-202 gains hard evidence: the orphan card promises those set-aside backups may be restorable later with their recovery code. For these 1.2 GB that is false and unfixable, and it is said to the customers most likely to read it. |
||
|
|
985f0ba63c |
Record the guards, the narrowing, and the two things I could not do
gates / gates (push) Successful in 30s
R-273's owed guards are both built and closed. R-291 records what CI stopped covering and why, so it can be widened deliberately rather than discovered. R-292 is new and was found by a test failing for the wrong reason: artifact_sha_invalid conflates "version missing", "registry unreachable" and "bad sha" into one message. v0.102.0 works around it by ORDERING -- the probes run first, so an unreachable registry is reported as unreachable -- but the message itself is untouched. CONTEXT gains the rule this session is about: a check and the policy it enforces must read the same number from the same place, or they drift and the drift looks like a defect in something else. Two corollaries, both of which cost something: a bounded check must print what it stopped covering on every run, and an unreadable policy is INCONCLUSIVE rather than unbounded. Stated in the report rather than glossed: Part 4 (finding receipts for the twelve downgraded claims) was NOT done and is a shortfall, not a decision -- splitting it would have produced exactly the half-checked green the exercise exists to prevent. Part 5 was droppable and dropped. The tag-push green is not re-proved tonight and is not claimed; the evidence offered is runs 190 and 216. |
||
|
|
36bcd12543 |
Deploy hub v0.102.0, and record that the package deleter is still not established
gates / gates (push) Successful in 21s
manifests/hub.yaml 0.101.0 -> 0.102.0. Image built and pushed, and verified served by the registry before the bump rather than after. R-287 corrected on two counts. My own sentence "no DELETE on the packages API appears in 48h of Gitea router logs" is WITHDRAWN: kubectl logs on the Gitea pod now returns nothing older than 2026-08-09 16:35 and contains zero api/packages lines even for requests I made myself, so the log never covered the window and its silence was never evidence. A second attempt to attribute the deletion also failed and the deleter remains NOT ESTABLISHED. Sources exhausted: no register row records a package prune (R-210 is WAITING-ON-OPERATOR, says "Nothing was deleted; this is a list, not an action", and concerns local Docker images); package_version has no soft-delete column so a deletion leaves no row; Gitea's action feed carries no package operation at all in the window; and a uniform newest-ten cap is not visible -- felhom-agent generic holds 10 but the container packages hold 19 each. It may be unestablishable from this side: Gitea keeps no package-deletion trail. |
||
|
|
6088afcbed |
Verify the standing picture against source: 12 downgrades, and the decay ran both ways
gates / gates (push) Successful in 21s
55 claims verified. Twelve moved, all downwards: walked 32 -> 20, built 5 -> 17. Register ceiling R-284 -> R-290. THE RULE DID NOT FIRE THE WAY IT WAS EXPECTED TO. Not one downgrade came from code moving under an old proof. All twelve came from step 1 of the same rule -- the cited evidence does not exist. Measured: of the 28 capability-map rows behind the page's claims, 8 carry a tests/ or audits/ path and 20 carry prose only. The green dots were drawn from rows that cite an argument, not a walk (R-290). The map, not the dataset, is what needs fixing -- it still says PROVEN-LIVE for all twelve. And once it ran backwards: fault.operator-email looked contradicted by R-182, but live source shows the backup_run_failures digest allowlisted, operator-only and templated, with recovery_unit_capture_failed now record-only. The claim is right and the REGISTER ROW is stale (R-289). The session went looking for stale proofs and found a stale defect. R-281 WITHDRAWN -- wrong in both directions, settled by the operator's mailbox. The tripwire DID fire (escrow_blob_served 10:19:41Z = 12:19 CEST) and false error-severity alarms fired too, for deliberate attended work (R-285). The measurement's cause is ESTABLISHED: the P7 query copied hub.db without hub.db-wal, and the signature is exact -- it reported "2 events all day, newest 00:30:07", and the rows at or before 00:30:07 number exactly 2. Timezone and wrong-key were tested and refuted. The control had been drawn from the same stale snapshot as the measurement, which is why it agreed (R-286). Part 4: NO WORKFLOW CHANGED, deliberately. The gate is not ref-sensitive -- it enumerates from the Gitea tags API, and both previous tag pushes passed. The red is TRUE: run 267 saw v0.120.0 downloadable, run 284 on the same commit saw 404. Who deleted the package is NOT established and is not guessed (R-287). The page is now generated from where-felhom-stands.yaml by scripts/render_stands.py: static, zero script tags, every moved status carrying a visible "changed, was X" chip. The React bundle -- whose content was gzip+base64 inside a JS module map -- is kept as a dated snapshot. scripts/check_stands.py gates the data and convicted 51 problems in my own first draft before the staged positive control ever ran. |
||
|
|
a199c492f4 |
Merge branch 'main' of https://gitea.dooplex.hu/admin/felhom.eu
gates / gates (push) Successful in 24s
|
||
|
|
fc4205f3b4 | html | ||
|
|
1d6f1c522d |
Rehearsal 2026-08-09 COMPLETE: data BYTE-IDENTICAL, journey needs a shell twice
gates / gates (push) Successful in 23s
The walk finished. All four planted files came back byte-identical out of snapshot 41c830db, including two Hungarian accented filenames verified as RAW NAME BYTES (NFC preserved) — the discriminator the Gate 0 positive control was built for, having been watched failing on an NFC->NFD rename that renders the same. Unlock 21s, restore 13.2s. It finished only because a terminal was available twice: - R-273 CLOSED. v0.128.0 was published as a package and never git-tagged, so every install died at 5/8. Tag pushed on operator instruction after an INDEPENDENT download proved the package sha equalled the vouched value; --resume then reached Day-0 SUCCESS in 3m49s on controller 0.210.0. The two guards that would stop the class recurring are still owed. - R-280 NEW, rank 1. A reinstalled box cannot re-attach its own data drive by any dashboard route: /api/disks/candidates returns empty because both lists are built from the UNCLAIMED-disk scan, and the drive is claimed precisely because it is also the backup target. Correct for "initialise", over-broad for "attach", which is non-destructive by definition. The restore page meanwhile says "Ez ket kattintas" and points at that empty list. Cleared by POSTing /mnt/sys_drive — an internal path no household could produce. Also new: R-281 the hub said NOTHING through the entire reinstall and the sealed-backup tripwire did not fire on a real unseal (positive control: 2 events all day fleet-wide); R-282 one code with three names and a mail pointing at a page the box does not show; R-283 hub reads "Claimed 18d ago" while the box serves its setup page; R-284 "almost full" over a 93%-free store. R-274 NARROWED by measurement rather than left as written: the resume path fetched the vouched golden correctly, because --resume skips the preflight that does local discovery. What survives is real — discovery is sort|tail -1 with no manifest comparison — but a FRESH install taking a stale golden is still not observed, and the row says so. Two of my own claims were refuted by test and are recorded as refuted, not quietly dropped: the leftover sudoers file is inert (sudo skips dotted names), and demo-hp's off-site tier was healthy all along. |
||
|
|
b1afbb8a4d |
Rehearsal 2026-08-09: the walk stops at P3 — R-273 blocks every install fleet-wide
gates / gates (push) Successful in 24s
P1 uninstall, P2 preflight, P3 install. The install FAILED at step 5/8 in 44s, and the two rank-1 findings are both on the setting-up path a tester's visit is made of. Eleven register rows minted (R-269..R-279); ceiling moves 268 -> 279. R-273 (RANK 1) — the hub vouches agent 0.128.0; that version was published as a Gitea PACKAGE but never git-tagged. Since R-183 the installer correctly pins its config fetches to raw/tag/v<vouched>, so every fresh install and every reinstall now 404s as root, mid-install. Measured: main 200, v0.127.0 200, v0.128.0 404. This is R-184 arriving; release-agent.sh:23 already documents the exact hazard. Existing boxes are fine (self-update takes the binary from the registry). NOT fixed here — publishing a release tag is outward-facing and the runbook says stop and report. One command unblocks it; it is in STATUS.md. R-272 (RANK 1) — Felhom's own uninstall leaves the condition that makes Felhom's own reinstall refuse. It installs dnsmasq at day-0, then on teardown removes the snippet and RESTARTS the daemon unconstrained (process start time lands inside the uninstall window), which grabs 0.0.0.0:53; the next preflight then refuses, and the message reads as though the owner's LAN DNS is at fault. R-274 — a local golden is adopted with no version and no sha check; the manifest vouch is consulted only on the fetch path. demo-hp's local copy is controller 0.192.0 against a vouched 0.210.0, and below the 0.200.0 where the recovery screen shipped. Not yet observed end-to-end (R-273 killed step 5 first). Also: R-275 orphaned credential backups + uid reuse, R-276 the wg tunnel outlives the uninstall, R-269/270/271 from the token rotation, R-277 three hub surfaces misreport a healthy off-site tier, R-278 demo-felhom six days unprotected, R-279 no operator-triggerable off-site run. Two hypotheses of mine were tested and REFUTED rather than shipped as findings: the leftover sudoers file is inert (sudo skips dotted filenames), and demo-hp's off-site tier was healthy all along - I had misread the hub and said so. STATUS.md records the three rulings §8.3 asked for, with the floor CORRECTED to its live value 0.200.0 and the count corrected to twenty. |
||
|
|
34646295dc |
Rehearsal 2026-08-09: pre-phase + Gate 0 recorded before the destructive walk
gates / gates (push) Successful in 29s
Venue demo-hp, operator-approved at STOP 1. Records the state that P1 destroys,
plus seven pre-walk findings, while they can still be checked against a live box.
R-268 CLOSED — the leaked per-guest local-API token is rotated and the rotation
is PROVEN in both directions (old refused, new accepted, channel up with a
positive observable). Rotating it surfaced three defects:
- an out-of-process rotation does NOT revoke the old token. The daemon serves
Lookup from a stale index and re-reads only on a MISS, so a superseded token
is a direct hit. Red-proved in a unit probe AND live on hardware; the shipped
RemintCoherence test passes only because it looks up the NEW token first.
- R-268's own recipe is incomplete: ensureLocalAPI returns early on a present
local_api block, so writing bootstrap.json is not enough — the controller
serves the old token from controller.yaml across restarts.
- the agent-channel alarm never closes: the UP branch does not notify from an
unseeded state, and the alarm's own remedy ("re-bootstrap") resets it.
Gate 0 complete: dataset planted in the Calibre library (coverage verified, not
assumed) with two Hungarian accented filenames; the comparator watched FAILING
three ways including an NFC->NFD rename that renders identically; off-site run
driven through the product's own button; restore point recorded by identity as
snapshot 41c830db, confirmed to carry all four files.
Also corrects the record: demo-hp's off-site tier is HEALTHY. Three hub surfaces
agreed it was absent and all three mislead — the panel showing 0 snapshots renders
the LOCAL tier, 162 KB rounds to 0.0 GB, and a two-day-old stuck event reads as
current. And the managed-update floor is live at 0.200.0, not 0.156.0.
|
||
|
|
56f8aa611c |
R-267 closed: 26.2s -> 5.4s cold / 0.14s warm, and two corrections to my own measurements
gates / gates (push) Successful in 40s
Registry pruned to the newest 10 per package on the operator's confirmed rule. 33 deletions, all HTTP 204; the live-vouched golden 0.210.0, agent 0.128.0 and floor 0.127.0 were asserted into the KEEP set BEFORE any DELETE was issued and verified still fetchable after. TWO CORRECTIONS TO WHAT I REPORTED EARLIER, both recorded rather than quietly dropped: 1. 'Only 50 generic versions exist' was NOT a count, it was a PAGE LIMIT. ?limit=1000 returns at most 50, and the 50 I measured was exactly the cap. Three older agent versions (0.81.0/0.80.0/0.79.0) only became visible after the first 30 deletions moved them onto page one. An unpaginated listing is not evidence of a total — this repo's own 'an empty listing is not evidence of emptiness' rule, walked into while measuring it. 2. The operator's 'reduce the number of artifacts' was the better call and my measurement said otherwise. I reported it helps sub-linearly and is not the lever. Measured after: trimming to 10+10 took the COLD load from 13.4s to 5.4s, a 2.5x improvement on exactly the path the memo cannot help, because the fan-out is per-version. drill-r50 runs agent 0.113.0, now deleted; flagged before deleting, disposable nested drill VM, only its re-download path is gone. |
||
|
|
efe9dfd15d |
R-267 CLOSED (26.2s -> 0.24s warm); R-268 filed: I printed a live token into a transcript
gates / gates (push) Successful in 12s
R-267 closed by hub v0.101.0. Measured after the 60s memo: cold 13.4s, warm 0.24-0.33s. The operator sees a quarter-second except at most once a minute. R-268 filed against myself. Setting up R-221's live drill, a one-liner meant to list bootstrap.json's KEYS printed the local_api object whole, including its token, for guest 9201. Reported rather than quietly rotated, because a secret reaching a transcript is a finding whatever its blast radius. Exposure assessed rather than assumed, and it is small: the token opens only the agent's per-guest local API on the island bridge between that host and that one guest, self-scoped to guest 9201, not routable from the LAN or internet, on a Tier-0 disposable box with no customer data. Using it already requires code execution there, at which point an attacker has more than the token. Rotation exists (TokenStore.Mint, last-write-wins per VMID) but must also rewrite the guest's bootstrap.json or the controller loses agent access — an operator-timed act, not a background one. The general fix is upstream: reading secret-bearing JSON should go through a helper that prints keys and never values, the discipline the golden bake already uses for the Gitea token. R-221 also recorded as PROVEN ON HARDWARE in STATUS. |
||
|
|
1c14b91d6f |
R-267: the Configuration page, measured — 26.2s to ~10s, and what is left
gates / gates (push) Successful in 22s
Reported as 'almost minutes'. The guess that it hashes artifacts on page load DOES NOT HOLD and the code already said so: Gitea stores the sha and the hub reads it as metadata. The cost was latency x count, fixed in three legs (hub v0.100.0-0.100.2), each found by refusing to accept a number that did not match the arithmetic. Measured 26.2s -> mean 9.85s over 8 samples (min 5.13, max 18.13). The remaining dominant cost is the package SEARCH, 0.20-3.8s per dropdown depending on load, which concurrency does not help; 16 concurrent file-metadata calls take 0.58s by comparison. EVERY NUMBER IS CONTAMINATED and the row says so: taken on DooPlex at load average 7-11 while this same session was building images, running two Go suites and baking a golden. The same search measured 3.8s in-cluster and 0.44s from the host ninety seconds later. Re-measure on an idle box. Both operator proposals answered on the measurement rather than deferred to: pruning artifacts helps sub-linearly (only 50 versions exist) and is worth doing for its own sake; storing the hash in the hub DB is NOT recommended, because Gitea is already the store and a copy would be a second source of truth the operator reads to confirm a vouch. The lever that would work — an in-memory cache with a TTL — is left OPEN because it trades dropdown freshness for speed, which is an operator decision. |
||
|
|
4a4a1e245a |
R-265 CI timeout + golden 0.210.0 baked; R-221/R-259/R-258 closed, R-266 minted, G-3 unblocked
gates / gates (push) Successful in 32s
Four defects of one family, all shipped today: something the box already knows, thrown away or drawn
as its opposite. Agent v0.128.0, controller v0.210.0. NO HUB CODE, no hub bump, no ArgoCD sync.
R-265 (this repo). timeout-minutes: 5 on the gates job — every honest run in the observed session
finished in 18-34s, so this is ~9x the slowest and far under whatever reaped run 264 at 834s with no
log. The alarm mail now carries Elapsed (start stamp via $GITHUB_ENV; an absent stamp prints
"unknown (no start stamp)", never a bogus 1.7-billion-second figure) and its "names itself in the run
log" sentence is qualified so it cannot mislead when there is no log.
⚠ THE UNKNOWN IS NOT CLOSED. Whether the if: failure() alarm fires for a REAPED job is still
unverified. The timeout makes the reap unreachable in practice; it does not answer what happens in
one. Demonstrating it means deliberately hanging a run on main, which would leave the branch red for
a parallel session. Said in the workflow comment, the changelog, R-265 and the report — none of them
claiming it is answered.
GOLDEN 0.210.0 baked, published, round-trip verified, NOT VOUCHED. The currency gate went red the
moment the controller was bumped — correct — and is closed by the bake, never --no-verify. No
--no-verify anywhere this session.
⚠ THE AGENT WAS NOT PUBLISHED UNTIL THIS SESSION CHECKED, AND IT MATTERED. R-221's fix is in the
AGENT, and a fresh install takes its agent from the Day-0 manifest. The binary had been hand-deployed
to felhom-pve and never published, so agent_version 0.128.0 was not selectable and a fresh install
would have received 0.127.0 — the golden would have carried the controller fixes and NOT the one the
headline defect needed. Caught by checking each Day-0 value was FETCHABLE rather than assuming.
Published from the live-deployed bytes, sha-verified across the hop first.
Registers. R-221, R-259, R-258, R-265 CLOSED. R-266 MINTED (READY): the failed root statfs still
travels to the hub as a 0-of-0 disk; ranked LOW because it is the quiet direction — it can only miss
a true alarm, never raise a false one — and it is now a two-repo wire change governed by G-1's gate.
Highest ID moved R-265 -> R-266.
CONTEXT S-39 rules the convention this project was missing: "we do not know" is never drawn as
"fine", and the codebase has ONE way of saying it — an explicit ...Known bool companion checked in
the template. ROADMAP G-3 was explicitly blocked on that decision and is unblocked; what remains
there is a survey-and-convert of existing sites, not the gate.
Capability map row 93 CHECKED and it was NOT claiming something untrue — it is about the operator
notification path. But its narrative ("the page you open to ask whether ONE app is backed up")
invites the wrong reading, and the adjacent thing WAS false until v0.210.0, so the row now records
that the two halves disagreed and only the operator half was true.
Six red-proofs across the two code repos, each with the mutation asserted applied. The one that
matters: Part 1 Scenario A FAILED against today's tree, with the intended message.
Part 1's operator-present live validation is OWED and is the session's STOP.
repo_gates --fast: all 8 OK.
|