99af997ab9e1e0c91c3cb48b38419e03d1795540
281 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
99af997ab9 |
R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163. |
||
|
|
4f875174fe |
Golden 0.226.1 baked, vouched, and the fleet floor raised — the debt is paid
gates / gates (push) Successful in 17s
golden_currency_gate.py had been CONVICTED four times today across three controller releases. One bake covers all three, and the gate went red -> green on the same command, which is its proof that it measures something real. THE THREE DECLARED BYPASSES ARE NOW HISTORICAL RATHER THAN STANDING. GOLDEN_VERSION 0.226.1 GOLDEN_SHA256 70ed8e9377dec22a9b493e55f222b0e25a49d7f3caec8c506e0412fd6baefe69 size 657 197 592 B baked gitea.dooplex.hu/admin/felhom-controller:0.226.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP, NOT THE BUILD LOG. The published bytes were downloaded back -- size and sha256 both identical to what the bake reported -- and ./etc/felhom-controller-image was read OUT of the downloaded archive: `felhom-controller:0.226.1`. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Acceptance markers counted, not eyeballed, each string captured from this run's own log rather than paraphrased from the runbook (two of the three the runbook named until R-233 could not match anything the script prints): docker OK (overlay2 = 1, including mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. The 404 pre-gate passed before the run, so nothing was overwritten. THE VOUCH IS A THREE-FIELD CHANGE AND ALL THREE WERE CHECKED: agent_version 0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash -- golden option 0.226.1 SELECTED, all four shas matching. THE FLOOR IS PROVEN ACTING, NOT MERELY SET. demo-felhom self-updated within 30 seconds: "[selfupdate] Post-update startup: update successful (0.225.0 -> 0.226.1)". Both demo machines now run 0.226.1 and only one of them was deployed to by hand. Token hygiene: copied file->file, read inside the VM by a runner script, never on a command line (systemctl show ... | grep -c -F token = 0). THE LEAK GREP ON THE COMMITTED LOG WAS PROVEN TO WORK BEFORE ITS 0 WAS BELIEVED -- a throwaway copy with the token appended grepped 1, was shredded, and only then was the real log's 0 taken as evidence. Teardown: build guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out (standing rule 5), VM powered off, disk reverted to virgin. R-242's OTHER half is untouched and still open: nothing gates the VOUCH itself. |
||
|
|
c8100aad6b |
Housekeeping + R-397/R-398 filed
gates / gates (push) Failing after 17s
Compresses the six rows closed today into CLOSED-ITEMS, keeping title, shipping
version, evidence paths and every sentence that states a RULE. Full original:
`git show
|
||
|
|
e027b5d999 |
Register + architecture for controller v0.226.0 (R-353/357/358/360/396), and R-395 fixed
gates / gates (push) Failing after 17s
Closes R-353, R-357, R-358 and R-360 with their shipping version and evidence path, and files two new rows. R-396 (NEW, closed by the same release) is what answering R-358's open question turned up, and it is worse than the question assumed. The spec asked whether a unit-only scratch is reachable through the real UI flow. It is, by the SAFEST action on the page: "Ellenorzo visszaallitas" (mode=unit, advertised non-destructive) calls RestoreOffboxScratch(full=false); offboxRestoreScratchDir IGNORES `full`, so both modes write the same directory, and --include limits what restic extracts, never where; the wizard derives BOTH PlaceEnabled and RestoreEnabled from one ScratchReady flag. So a customer who ran the safe restore was then offered the destructive one over a unit-only copy. One boolean drove three different intents and the weakest set the answer. R-395 (filed by the spec) is fixed in this commit, not just recorded. STATUS.md said golden 0.223.0 / floor 0.222.0 in one block and demo-hp 0.219.0 / floor 0.218.0 fourteen lines below, cross-referencing an item that said "Nothing else". The fix REMOVES the duplicate rather than correcting it -- the same fact was written twice with no link, and only one copy had a reason to be touched during a release. "What works" now points at the item above instead of restating a version. 07-backup-architecture: four rows added to the 10.2 gap register plus R-396. Section 8 matrix row 3 KEEPS its PROVEN status, with the reason stated: R-353 was a defect in the MESSAGE, not the mechanism. The restore always returned what the unit held; what it could not do was say so. A status that measures whether data comes back must not move because a status line was wrong. 00-capability-map: one new row, and it splits what is claimed. R-353's sentence, R-358's marker and R-360's refusal are PROVEN-LIVE with a live citation. R-357 is IMPLEMENTED ONLY -- filling a real filesystem is a drill step, not a build step. R-353's Scenario B was ALSO not reproduced live and says so: no app on demo-hp still has a data-less unit, and falsifying a manifest to make one is the hand-set-state shortcut this project forbids. This push used `git push --no-verify`. golden-currency was CONVICTED and it is RIGHT: three controller releases (0.224.0, 0.225.0, 0.226.0) and the golden still carries 0.223.0. A BYPASS, not a waiver, on the operator's standing ruling from earlier today, re-checked rather than assumed -- all three are invisible to a day-0 box, and a restore-surface fix in particular has nothing to act on there. The ground expires the moment a release changes first-boot behaviour. Tracked on R-242; ONE bake carrying 0.226.0 covers all three. |
||
|
|
36f8630020 |
R-341 check taken, and the golden-currency bypass declared
gates / gates (push) Failing after 16s
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the difference is stated rather than blurred. FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z here -- R-346's trap). Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0, CLOSE-WAIT 0. The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade, and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. Row closed as moot. What it DOES establish is worth more than the original question: twelve days after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits at baseline with zero established connections. R-336's ~323-day runway concern retires with it. BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and the newest golden bake carries 0.223.0, so a machine installed right now gets neither. The gate is RIGHT. This push therefore uses `git push --no-verify`, declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242. A BYPASS, not a waiver: the gate offers a waiver only for a release that DELIBERATELY needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box -- R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub display over backups a new box has not taken yet -- and both arrive by self-update. That ground is recorded because it is what to re-check: it does NOT extend to a release changing first-boot behaviour. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1, three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap is now two releases wide rather than one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
c30430c530 |
skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394. |
||
|
|
ebdc04601d |
docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or coarse, and why the default is coarse. CONTEXT records two rulings: the grain is allow-listed rather than inferred from the payload, with crossdrive_failed as the proof that a payload rule would have been wrong; and a finding recorded only in REPORT.md has a lifetime of one session. R-389 closed and compressed, keeping its rules and naming the commit whose git show returns the full text. R-390 and R-391 left open. REPORT.md is gate 11's first real subject and passes: six observations, two FILED, four NOT-A-FINDING with their reasons. Three of those declarations are things a tidier report would have omitted - the gate's own spec would have passed the item it was built to catch, the burst has no ceiling, and ArgoCD said "successfully rolled out" while still running the old image. STATUS carries forward the one thing outstanding: the controller floor still reads 0.222.0 while the golden reads 0.223.0. |
||
|
|
2fc4a15fa3 |
R-389: key the operator cooldown per app for app_start_failed; gate 11 makes an unfiled observation refuse the push
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and none of them names an app, so every app going down inside the same hour collapsed onto one key and only the first was mailed. Measured on demo-hp: bookstack sent 09:27:51, privatebin suppressed 09:31:51 under key=demo-hp:app_start_failed. cooldownStackSuffix is the third sibling of cooldownTierSuffix and cooldownRunSuffix, and separate for the reason the second one's docstring already gives: the existing two keep byte-identical semantics for every type that uses them. It is ALLOW-LISTED to app_start_failed and takes the event type as well as the details, unlike its siblings, and that asymmetry is the safety property. The backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app - and crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a DIFFERENT struct, so a payload-shape rule would have split it silently. The hour itself does not change. Gate 11 refuses a push whose REPORT.md carries an observation with neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a passing mention of some other R-number: the lost item cited R-182 as an analogy, so "cites a register row" would have passed the very item the gate exists to catch. That discrepancy with the spec is recorded in the gate's docstring. Registered here and in the controller and agent runners. NOT in the catalog runner - it has no shared-gate mechanism and appends --all to every gate; filed as R-391 rather than left as a sentence, which is this session's lesson. PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording that invited the gap, and it now names the markers and points at the gate. R-390 filed for the golden-bake runbook's missing `pveam update`. Hub tests 709 -> 716. |
||
|
|
f751aea4e3 |
R-389: file the cooldown-grain finding that was never filed
gates / gates (push) Successful in 16s
Only the first broken app per hour reaches the operator. The cooldown key is customerID:eventType plus the tier and run suffixes, and neither reads an app name, so every app that goes down inside the same hour collapses onto one key. Measured on demo-hp 2026-08-23: bookstack sent at 09:27:51, privatebin four minutes later logged `suppressed - operator cooldown 1h, key=demo-hp:app_start_failed`. Filed FIRST, before any code, for two reasons. It should have existed since yesterday and did not - it lived in a REPORT.md observations paragraph and nowhere else, which is R-341's shape one surface over. And the gate this session adds refuses a push whose report carries an observation with no row behind it, so the row has to precede the gate or the gate refuses its own commit. |
||
|
|
2f7c9a6ce5 |
docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B. |
||
|
|
55274d5ef3 |
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act. |
||
|
|
1eb64bec51 |
R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
gates / gates (push) Successful in 17s
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an invariant the code did not have, for four months - and a [DESIGN] on the db_dumps decision INCLUDING the trap it created: a stable list lets the already-current early return fire, so per-capture housekeeping must sit above it. 00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which IsDownState excludes. Measured on the shipped build with the scans demonstrably running over it. No suppression was built and no row opened. R-383: the double-failure message names an undo copy that is not there - R-361's own class, one surface over, observed on both 0.220.2 and 0.221.1. R-384: an app whose database has died reads unhealthy and raises no alarm. R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes. Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate blocked this push and that block is not circular, so it was satisfied rather than bypassed - no --no-verify anywhere in this session. |
||
|
|
a8caa0fdde |
R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step. |
||
|
|
4e488321bf |
DRILL R-356b: the off-site restore for a driveless app that HAS a database
gates / gates (push) Successful in 16s
A drill, not an implementation. No code, no version bump, no CHANGELOG entry. Ten of the forty driveless apps carry a database; I re-measured that count and got 10. For those ten the restore is a five-leg operation that never ran at all until this week, because R-356 refused before any of it started. Walked end to end on demo-hp for both engines - docmost (Postgres 16) and bookstack (MariaDB 12.3) - each deployed for the drill, planted through the app's own interface, destroyed for real, restored through the endpoint the UI posts to. Q1 does it complete: YES. All five legs ran and succeeded, 32s / 25s. Accented names byte-identical both directions. Q2 which leg won: the SQL DUMP. Three-way discriminator returned the altered dump's value. This confirms R-164's F17 ordering on the OFF-SITE path; R-164 only ever cited the local one. Scratch-only mutation; store proved unmutated. Q3 does a failure tell the truth: partly, and two defects. Filed R-379 (HIGH, the undo copy is valid, named, and unappliable by any product action - proven by applying it by hand on both engines), R-380 (HIGH, a failed MariaDB replay leaves a partial database behind an app reporting healthy, where Postgres crash-loops visibly), R-381 (MEDIUM, the failure message pastes engine stderr including customer table rows into the Hungarian surface), R-382 (LOW, the summary log omits the volume count it already has). H1, H2 and H4 did NOT fire and that is recorded. H3 fired in a shape nobody predicted: not a quiet success, but a loud error over a silent inconsistency. R-361 reproduced independently on a second app. restic check: no errors, 29 snapshots. A flaw in the drill's own planting - a double-escaped accented title - was caught by reading stored bytes as hex, recorded, and re-measured in Phase 1b. Register 325236 -> 330683 bytes. Nothing dropped. Teardown: two apps retained with reason, no pvesm before-snapshot taken (said plainly), no hub-side record created. |
||
|
|
c297b9f85e |
R-356 docs: correct R-107 in the architecture, record the design, refresh STATUS, compress the register
gates / gates (push) Failing after 17s
07-backup-architecture.md: three places said no offsite action unpacks the named-volume tars. R-107 closed in controller v0.218.0; all three corrected with a dated [FACT], the old sentence kept in the past tense. R-102 is NOT closed and the correction says so explicitly. New [DESIGN] paragraph in 6.3: the restore destination is resolved by the same rule as the capture destination, and the wrong-disk refusal applies to apps that have a drive to get wrong. Carries the 13/40 measurement. STATUS.md was internally contradictory - nothing waiting, and one decision waiting, for something the same page recorded as shipped. 218 -> 102 lines; the deciding section now says what happens if nothing is done. R-356 compressed into CLOSED-ITEMS.md; OPEN-ITEMS 327109 -> 325236 bytes. Drill record and 16 evidence files for the live walk on demo-hp. |
||
|
|
ef6ac6fe74 |
One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)
gates / gates (push) Successful in 16s
Records and process only. No machine contacted. ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days. 15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header: does the item assert something about the shipped product a reader could check and find false? scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted open roadmap-only row is convicted by name, removing it passes with the file byte-identical, and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring. The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as done because my regex matched the whole row where the body contains "shipped" - the gate matches the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict. HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%). Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the commit whose git show returns the full original text. Rule-sentences are kept verbatim under "Reasoning kept" rather than judged entry by entry - 25 carry one. CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric. The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1 of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376). PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state the register's size before and after. Ceiling R-375 -> R-378. |
||
|
|
091a4b7444 |
Correct the placement mis-framing, and file what we wrote down and never filed (R-368..R-375)
gates / gates (push) Successful in 16s
Documentation and survey only. No code, no machine contacted. THE CORRECTION. The 40 catalogue templates without a configurable path are not missing a choice: 01-topology-and-trust.md:150-152 classes each volume hot (DB/config/cache -> fast storage, ENFORCED) or bulk (media/files), and the 40 are all-hot apps. The deploy page has been saying so to the customer all along (deploy.html:624-625). SPEC-app-data-placement and R-352 are corrected in place with the framing MARKED, not deleted; every measurement stands. R-356 was re-checked and survives, strengthened - an absent HDD_PATH is the normal state, so reading it as "not installed" misreads a correct configuration. The disk claim, precisely: since R-165 there is ONE guest data volume with two binds, not two volumes (build-golden.sh:29-40, 99). A physical-disk failure losing data and first-tier copy together is REAL and is what the other tiers exist for. A full data volume stopping the OS is NOT real and was the overstated one. THE SWEEP. 113 survey-class documents examined, 14 statements of "not filed", 2 already filed. Its positive control convicted the sweep itself twice before it convicted the corpus - markdown bold broke the strongest pattern, and the reporter re-searched a truncated line - both false zeros of the exact class being hunted, and together worth 2 of the 14. THE HEADLINE. The gap the 2026-08-21 drill rediscovered WAS filed - as R-107, ROADMAP.md:122, M/READY, 2026-07-28 - and is absent from OPEN-ITEMS.md, which calls itself the single source of truth. OPEN-ITEMS and that rule both landed 2026-07-27; R-107 went to ROADMAP alone the day after. 72 ids live only in ROADMAP, 29 not done, some of them findings. Filed as R-369 (HIGH). Five more still-open gaps filed with their ages: R-371 (17d), R-372 (38d, the oldest), R-373 (20d), R-374 (14d), R-375 (4d). R-368 corrects Part 4: the storage default IS applied at deploy time via deploy.html:612 - the earlier "the deploy route never reads it" came from grepping Go and never the templates. R-370 records the process failure and is closed by the template change. PROMPT-TEMPLATE gains the two rules it lacked: name the architecture document for the area and say what it says (with a file->area map and the test "is this something we chose?"), and an enumerated gap becomes a register row in the same session - a ROADMAP row alone does not count. Ceiling R-367 -> R-375. |
||
|
|
877fcd2a38 |
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367. |
||
|
|
7064596c2e |
DRILL closeout: the scheduled cycle agrees, the abandonment sweep watched firing, R-366
gates / gates (push) Successful in 16s
The box's own 02:30 / 03:30 / 04:15 cycle ran unattended and agrees with the manual one on every measure. The one that matters: R-355 is not an artefact of manual triggering — the scheduled run again wrote paperless-ngx's PostgreSQL dump into a directory for a stack that does not exist, and again left the app's own unit recording db_dumps: null. Part 4.2's terminal deletion has now been observed. At 05:10 the sweep removed exactly the recorded set-aside store and left the live repository untouched; at 05:13 the hub dropped the sealed package that protected it and said so (event 3025). Both halves went together, three minutes apart, and the controller cleared its own state. demo-felhom's two preserved fixtures were verified untouched throughout. R-366 (HIGH) filed, found incidentally: the 21 August reinstall orphaned demo-hp's PBS whole-guest archives as well as its restic repo, so a rebuilt box loses BOTH off-premises tiers at once. The restore-test caught it and named the key mismatch precisely; it is merely called "a failed restore test" rather than "your older backups are unreadable". demo-hp is left HEALTHY, not broken. Ceiling R-353 -> R-366. |
||
|
|
f5a4fceeeb |
DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
gates / gates (push) Successful in 16s
Diagnostic only — no code changed, no version bumped, nothing deployed.
The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.
Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
R-354 off-site restore never replays volume dumps
R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
has no DB dump, no safety dump is taken, and the customer is told it has none
R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
"is not installed", with a remedy those apps make impossible
R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.
Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
|
||
|
|
67356c9e5c |
register: R-351 CLOSED, R-352 partly, R-353 OPEN (next session's first item) + placement spec
R-351 - the restore never read back where the backup said the data lived, and a second press
started a second restore. Both shipped in controller v0.217.0.
R-352 - four measured untruths about where an app's data goes:
1. 40 of 53 catalogue templates declare no data path (13 declare env_var: HDD_PATH)
2. GetDefaultStoragePath() has three non-test callers and NONE of them places data; its
field comment "new apps use this by default" has never been true
3. the first-tier backup follows the data onto the same disk (GetAppDrivePath ->
systemDataPath) - the posture Tier 2 refuses outright at tier2.go:329
4. "1 alkalmazas hasznalja" counts only Env["HDD_PATH"] == path, so it can never include
the 40-class; it means "1 of the apps that CAN use a drive does"
Visibility shipped tonight; PLACEMENT IS OPEN and is the operator's ruling. An earlier
recommendation to refuse deployment until a drive is registered was WITHDRAWN - it assumed
the customer had failed to choose, and they had no choice to make.
R-353 - a restore reported success having returned configuration and no data. OpenGist's unit
holds manifest.json + compose/ and nothing else (volume_dumps: None, db_dumps: None); the
off-site snapshot was 182.3 KB; the outcome said only "completed in 8.666896042s". Ranked as
the NEXT SESSION'S FIRST ITEM. Compounding and recorded as UNKNOWN rather than fine: whether
the 40-class reaches the off-site tier at all has not been observed - runVolumeDumps covers
them on paper, but no nightly dump run had happened on a one-hour-old box.
New: documentation/backlog/SPEC-app-data-placement-2026-08-21.md - specification only, nothing
implemented, listing the five points a placement ruling must settle. Records that the OpenGist
instance meant to be left as evidence was removed by someone between 16:57 and 17:02 UTC (not
by this session); its unit and manifest survive, and privatebin is now a live specimen.
Ceiling moved R-350 -> R-353.
|
||
|
|
38ca4cf6f1 |
R-344 CLOSED: delivery was the last thing holding it open, and R-347 closed that
gates / gates (push) Successful in 14s
0.130.0 is tagged, published and vouched, so a fresh install gets the fixed agent. Both boxes reinstalled from the downloaded artifact. Closing evidence is the positive observable: ep0 at fd 17 / ESTAB 0 / CLOSE-WAIT 0, its t0 baseline, returning there between poll cycles with both agents still demonstrably polling. ep0 was read-only throughout and its proxy PID never changed. |
||
|
|
910fd91124 |
agent 0.130.0 published and vouched; R-347 closed, R-349 + R-350 filed
gates / gates (push) Successful in 14s
Released via scripts/release-agent.sh: tag v0.130.0 at 7569f34, sha256
a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3,
verified by independent download and reproducible byte for byte with
-trimpath -buildvcs=false.
Vouched agent 0.129.0 -> 0.130.0 in the Day-0 manifest. Only the agent
fields changed: min_agent stays 0.129.0 because it states what the GOLDEN
CONTROLLER requires, and raising it would have HELD the floor for every
box below 0.130.0. Global floor untouched at 0.216.0 -- and on hub
v0.106.0 it is a separate form with its own action, so publish-train
rule 2's hazard no longer exists in the shape its incident describes.
No --no-verify: the CHANGELOG heading was flipped only after the tag and
package existed, so release-complete passes on the real artifact.
R-349: the fleet was running a DIFFERENT binary under the same version
name -- the proof deploy was a hand build, the release is -trimpath.
Self-update could never have corrected it, because every version check
compares the string. Both boxes reinstalled from the downloaded package.
The proper fix exists in miniature as wrapper_sha256 and was never
extended to the agent's own binary.
R-350: I printed the hub password into the session transcript via
curl -w '%{redirect_url}' -- the hub answers 303 and curl re-attaches the
credential. Not in git, not in any committed file, not in the evidence
directory. Rotation is the operator's call.
ep0 closes at fd 17, ESTAB 0, CLOSE-WAIT 0 -- its t0 baseline -- and was
read-only for this entire arc.
|
||
|
|
57dd62b097 |
R-344 fixed and proven on both boxes: ep0 is back to fd 17 from 415
gates / gates (push) Successful in 14s
P1, outcome (i) in one second: replacing the agent on demo-hp released exactly its 199 established connections (ep0 fd 415 -> 216). CLOSE-WAIT stayed 0, so outcome (ii) does not exist and gets no row -- ep0 reaps on peer FIN correctly, and the 543 CLOSE-WAIT at the 08-18 wedge has another explanation. P2, 1.03 h (operator closed the >=4 h window early, so no daily rate is extrapolated): control +4, fixed +0, with each box making exactly 4 /snapshots and 4 /version calls. Same cadence, same work: 4 cycles -> 4 leaks vs 4 cycles -> 0. The fixed box's cycles are in ep0's log, so the zero is the fix and not a stopped agent. P3: the second box took ep0 from 220 to 17 fd in under two seconds. 17 is precisely the t0 baseline of 2026-08-18 09:51:22Z. Corrects a claim this session made earlier the same day: the accumulated descriptors did NOT need an ep0 proxy restart. They were held on both sides. ep0 was read-only throughout; its PID never changed. R-344 updated and left OPEN (unpublished is not delivered). R-336 re-scoped -- its old next-step would have fixed nothing while looking like a failed fix, and it is now a scaling row (~25 req/s at fifty customers). R-347 filed for the delivery gap (Viktor decides). R-348 filed: an agent restart blanks the reported backup list for ~18 h and the Store comment calls it unaffected -- blinds no alarm, checked not assumed. |
||
|
|
19672e685e |
SPIKE ep0 connections: the leak is felhom-agent's, not the poll rate
gates / gates (push) Successful in 14s
R-341's first dated check, taken at +46.2 h: fd 17 -> 405 over 166,251 s
= 201.6/day. Pre-registered range was 370-450; observed 388. UNCHANGED,
as predicted. CLOSE-WAIT is 0 -- absent entirely, not merely flat.
Q1: exactly two peers, 194 each, no third party.
Q2: outcome (a). ep0 388 = 194 + 194 on the boxes, twice, and the four
new sockets carry the same source ports on both sides. 0 closed in 31 min.
The finding: all 388 are held by felhom-agent. pvestatd and
proxmox-backup-client made 162,404 requests and leaked zero. Mechanism is
a per-cycle http.Transport with a zero-value IdleConnTimeout that nothing
ever closes (internal/pbs/client.go:56, main.go:1486). R-336's premise
does not survive this -- cutting the poll rate would have fixed nothing.
Q3 NOT measured: Phase C held at STOP 1, prediction pre-registered first.
Part 0 captures the due-checks gate's first conviction on a real overdue
date (rc=1, names R-341, sole failure among 10 gates). Row cleared at
Part 4, after the result was recorded in R-341, not to make a push work.
New: R-344 (the transport leak), R-345 (hub/Makefile pushes :latest),
R-346 (ActiveEnterTimestamp reads 5h56m early -- NRestarts is still 0).
|
||
|
|
ab2262c91c |
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
|
||
|
|
0a5e9b14dc |
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10. |
||
|
|
f267bc047f |
R-334 CLOSED, quoting the green CI run id
gates / gates (push) Successful in 14s
Golden 0.216.0 baked, published and vouched; the row is closed with CI run
353 (head_sha
|
||
|
|
3e50902a98 |
RUNBOOK ep0: PBS 4.2.2-1 -> 4.2.5-1, slope unchanged as predicted (R-341)
gates / gates (push) Failing after 14s
Both STOPs cleared by the operator. No code changed; documentation only.
STEP 3 (the run's primary deliverable): the full changelog range 4.2.2-1 ->
4.2.5-1 was read (128 lines, all three entries) and swept for
connection-handling vocabulary. Exactly one keyword hit, a false positive
("S3 ... honor the node's proxy settings" = HTTP proxy config for S3, not the
PBS proxy daemon). 4.2.5-1 is a manifest-hardening security release; 4.2.4-1
is S3 rate limits and a locking cache; 4.2.3-1 is UI/LDAP/tape. NOTHING
addresses descriptor lifetime or connection reaping. Recommendation was: do
not upgrade for this reason.
STOP 1: operator ruled to upgrade anyway for rehearsal value. Recorded as a
practice run, not a fix -- and the interpretation was fixed IN WRITING BEFORE
any numbers existed (stop1-ruling.txt): unchanged = expected; changed =
surprise. Neither outcome could then be rationalised into a success.
STOP 2: Hetzner snapshot 421440873, Available. Documented that it covers
/dev/sda ONLY -- /mnt/pbs-datastore is a separate Volume and is NOT in it, so
it is a software rollback and not a backup of the backup data.
UPGRADE: simulated first (0 to remove), then installed 09:51:00->09:51:06Z,
exit 0. Verified: 4.2.5-1 installed, both daemons active, effective open
files still 65536 (the drop-in survived the new package), Recv-Q 0, loopback
200, 200 from BOTH boxes over the tunnel with felhom-pbs active, and the hub
gauge refreshed post-upgrade at 11:59:31.
SLOPE: before +4 fd/1885 s = 183/day; after +5 fd/1919 s = 225/day. NOT
distinguishable -- one descriptor apart, Poisson +/-2 on such counts. The
higher after-figure is noise, not a regression and not an improvement. 30
minutes cannot settle it; R-341 files the +24 h and +7 d checks.
CORRECTIONS to this morning's own report, both published rather than quietly
fixed:
- the "~85/day, ~2 years of runway" figures were WRONG. They came from a
single 17-minute window with a delta of ONE descriptor. Real rate is
183-200/day over two independent windows; runway ~357 days, not 2 years.
- the leak was attributed to CLOSE-WAIT. It is mostly ESTAB: CLOSE-WAIT held
flat at 1 while ESTAB grew 45->49, and at the wedge it was 1011 ESTAB vs
543 CLOSE-WAIT. R-336's fix must target unreaped connections.
- "proxmox-backup-api" reported inactive during verification; that unit does
not exist. Bad query, not a fault, written down because it looked like one.
R-336 stays open: even a fixed leak would not make ~85k requests/day to a
weekly-write DR endpoint correct.
golden-currency still convicts (inherited R-334, controller 0.216.0 vs golden
0.214.0, untouched by this run), so this push is --no-verify per
.claude/rules/gates.md.
|
||
|
|
435e044cf1 |
INCIDENT/R-336: the leak is measured live, not assumed
gates / gates (push) Failing after 15s
Post-restart baseline on ep0: 19 fds at 16m43s (from 18), 1 CLOSE-WAIT, Recv-Q 0. One descriptor per ~17 min is ~85/day, which agrees with the ~73/day implied independently by the failure itself (1016 sockets over 14 days of uptime). Two estimates of the same slope agreeing turns "the ceiling raise is mitigation, not a cure" from a plausible claim into a measured one, and puts the next ceiling at ~2 years instead of a fortnight. Recorded because standing rule 3 asks for a positive observable: this is it, and it fired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
ebfd0967c1 |
INCIDENT + registers: ep0's PBS proxy served nobody for 9.5h (R-336..R-338)
gates / gates (push) Failing after 12s
Two whole_guest_backup_failed alerts at 04:30 and 04:32 CEST were one incident, and not on either customer box: ep0's proxmox-backup-proxy was active, holding its listening socket, and accepting nothing. Root cause: accept() returning EMFILE. The process held exactly 1024 fds -- its systemd-default soft RLIMIT_NOFILE -- of which 1016 were sockets and 547 connections sat in CLOSE-WAIT. The 1024-deep accept backlog had overflowed (Recv-Q 1025), so every client timed out. It was wedged from its own loopback too, which is what moved this from a network problem to a process problem. Fed by ~85k requests/day (a flat 3,538/hour) against an endpoint written to weekly, that leak reached the ceiling in 14 days of uptime. Fix: LimitNOFILE=65536 drop-ins for both PBS units, restart, verified from both boxes (200 in ~0.1s, felhom-pbs active), then re-drove the missed backups through the product path -- POST /backup?target=felhom-pbs on each agent's local API, not a hand-run vzdump. demo-felhom ct/9201/2026-08-18T03:57:43Z 4.10 GB 36.4s demo-hp ct/9201/2026-08-18T03:58:43Z 4.29 GB 41.5s Both host reports now carry felhom-pbs success=true, so the hub is green on the evidence rather than on a restart having been performed. No data lost, no backup skipped: the daily local tier was never affected and the PBS tier is weekly, so the window cost exactly one attempt. Evidence copied off ep0 BEFORE the restart, per standing rule 5. Filed: R-336 (the ~1 req/s poll rate is the real defect; the raised ceiling is mitigation, not a cure), R-337 (a status endpoint that trailed its own artifact by minutes then caught up -- WATCHING, downgraded from the defect I first wrote, because it self-corrected), R-338 (demo-hp is not on the R-50 island at all and nodes.md says it is; its local API is bound to the customer LAN). R-334 updated: still open, now one version wider (controller 0.216.0 vs golden 0.214.0). golden-currency is the only failing gate and is inherited -- it reads files this session did not touch -- so this push used --no-verify, stated per .claude/rules/gates.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016p1PTCzb8rF5G9Aa1qhpBN |
||
|
|
ea16a21bff |
docs(register): narrow R-332 — restart persistence is now proven live
gates / gates (push) Failing after 14s
The 0.215.0->0.216.0 redeploy replaced the container and the state file came back with the previous version's changed_at, so the new container loaded the pre-restart record rather than re-baselining. The verdict path itself, and the already-alerted-disk restart case, remain unproven. |
||
|
|
fa4748d4dd |
docs(register): R-335 — one physical disk walked twice per run, sustained against itself
gates / gates (push) Failing after 16s
Found on live hardware after the v0.215.0 deploy by reading the check's evaluated count against its own persisted state file. Closed in v0.216.0. |
||
|
|
767960bb11 |
docs: disk-health phase 1 — capability map, roadmap arc, register rows R-328..R-333
gates / gates (push) Successful in 14s
- capability-map: the disk-failure scenario no longer says a failing disk has never been seen. Healthy path + delivery + the severity wire stay PROVEN-LIVE; the new Hiba-from-counters path is IMPLEMENTED and explicitly NOT proven-live (R-332), because it has only ever run against the fixture's values. - ROADMAP R-73: phase 1 shipped; the genuinely hub-side half splits into R-330 (phase 2, a declared wire change under G-1) and R-331 (phase 3, growth-rate detection and retiring the static 64). Its premise 'no demo hardware exposes real SMART' is retired — a real failing drive is now committed as a fixture. - register: R-328 (the severity drop, CLOSED and proven live side by side), R-329 (app_start_failed has the same defect, needs a decision first), R-330/R-331 (phases 2 and 3), R-332 (the Fail path has never fired on real hardware, WATCHING), R-333 (NVMe temperature bands measured 2 degrees from tripping on a healthy drive; and the agent's smartctl has no -n standby). |
||
|
|
b03a105375 |
hub v0.105.0: the third name, a machine told to be quiet, and a guard for the hub's own words
gates / gates (push) Successful in 17s
Hub only. No controller change, no agent change, no wire change — nothing to bake. demo-hp untouched: the operator is re-deploying it this evening. R-323 — the five-word phrase is „Tulajdonosi jelmondat". It was „Visszaállító jelszó": one word from the name retired last week, and false besides — it restores nothing, it proves the account owns the box being bound. Five sites, all in the hub; felhom-controller and felhom-agent carry the name nowhere, so no halt and no bake. Both suggested names were rejected with reasons: „Fiókjelszó" would collide with the dashboard login (a DIFFERENT real secret), and „Összekötési jelszó" would leave the two factors on this page separated only by kód-versus-jelszó — the exact shape being removed, since the other factor is the „Párosító kód". The chosen name differs on both axes, stem and noun. Naming only; the acceptance pin drives the real handler. R-324 — the hub's customer copy is under a guard for the first time. Retired names banned across all 95 hub files; retrieval stems registered in four declared customer surfaces. The selftest found a defect in its own instrument on the first run. One shared vocabulary in scripts/, drift-checked into the controller gate rather than copied (R-325 removes the scaffold). R-321 — a machine we told to be quiet is no longer reported as dead, and it was two doors, not one: because the state is RECORDED rather than deleted, the morning deadline check can skip it too. A deleted state returns "", which is not "down" — R-195's shape returning through a second door. The clock runs from the report the hub can see, so re-enabling starts it there and emits no recovery for an outage that never happened. Three red-proofs; the one that matters showed a genuinely dead machine sitting at "disabled" when the suppression was made unconditional. R-326 — "which claims are unproven" is answerable by a command now. The nine I have been repeating was the count of claims the 9 August pass DOWNGRADED, not the count of unproven ones. The real figures: 55 claims, 23 walked, 32 not — and only 6 of those 32 cite evidence. Its first run found a stale claim (R-327). |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
fc737b0fc0 |
installer v1.28.0: the removal genuinely reverses the installation (R-316)
gates / gates (push) Successful in 13s
v1.27.0's fix worked exactly once per machine. Measured on drill-r50 from virgin, on the PUBLISHED v1.27.0, before anything was changed: cycle 1 recorded 'no' and freed :53; cycle 2 recorded 'yes' and left dnsmasq running on 0.0.0.0:53; cycle 3 refused, exit 1. Every box already in the field is at cycle 2, and a reinstall onto a machine that has had Felhom is cycle 2 by definition. Why cycle 2 says yes: the preflight's ownership question is dpkg-query package presence and nothing else - not the absence of a record. Stopping the unit and leaving the package made our own package read as the household's one cycle later. Now the uninstall removes the package when the record says we installed it. Order unchanged and load-bearing: read the record, act, then delete the state file that holds it. TWO packages are recorded, because dnsmasq ships the unit and dnsmasq-base ships /usr/sbin/dnsmasq, and each is taken back only if we added it. The dependency check is a SIMULATION, not a guess: apt-get -s purge is asked what it would remove and the purge proceeds only if that set is a subset of ours; otherwise stop+disable, naming the package that blocked it. Never interactive, never fatal, and the success is re-queried rather than read off an exit code. Watched: three fixed cycles -> install 3 PASSES; a household resolver untouched; a dependent package not purged and named; no record -> untouched with the command named. Red-proofs with the mutation asserted applied: remove the purge -> cycle 3 refuses in those exact words; remove the ownership check -> a household resolver is purged; infer ownership -> the guess is taken. Also: R-317 (the agent stats a path dnsmasq-base owns to decide whether to install dnsmasq - pre-existing, now reachable), R-318 (no honest ownership marker exists for existing boxes; the preflight message is the mechanism), and the status page's decisions section rewritten to say what each decision costs and what doing nothing selects. |
||
|
|
d102ca5767 |
Vouch delivered fleet-wide; the MinAgent hold observed firing and releasing for the first time
gates / gates (push) Successful in 20s
Operator vouched golden 0.214.0 / agent 0.129.0 / min agent 0.129.0 and raised the floor to 0.214.0. Both artifact shas in hub_settings match the bake and the release byte-for-byte. With the floor at 0.214.0 and demo-hp still on agent 0.128.0, the hub HELD the controller floor - 'agent 0.128.0 < MinAgent 0.129.0 (controller floor withheld)' - and released it 6 seconds after the agent was brought up. That is the R-216 guard doing exactly what it exists for, seen firing for the first time, and it is the argument for declaring MinAgent in the CHANGELOG header. Both demo boxes verified on the boxes: controller 0.214.0, agent 0.129.0, healthy. |
||
|
|
684cd2eb11 |
R-265 third sighting: the jobs API names the failing step when the log 404s, and a re-run disambiguates
gates / gates (push) Successful in 15s
|
||
|
|
4906aeb3f9 |
R-311 proven live (HTTP 422 on hardware); R-308 WITHDRAWN — my quoting bug, not a stale credential
gates / gates (push) Successful in 19s
The live test read as a FAILURE for twenty minutes because I stripped only double quotes from a credentials value wrapped in SINGLE ones, sending a literal ' as part of the recovery code. Correctly unquoted: the old code returns 422 with opens_retained=true and the supersession date; a wrong code still returns 400. The same bug produced the R-308 finding in the previous report. The dashboard password is fine - HTTP 302 with a session cookie on the first try. Third time this project has produced a wrong 'the credential is stale' verdict from that one trap. |
||
|
|
6362bb6cb6 |
hub v0.103.0 — a host can read the packages we kept for it (R-311)
ListSupersededEscrow had zero production callers for nineteen days. It is the only reader of a retained identity_blob, so the retention shipped in v0.93.0 was material the product could not reach - proven on the fixture 2026-08-12, where a code that opens a retained package was answered as a code that opened nothing. New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row GET, same recovery-mode gate, same audit event written BEFORE the bytes leave, capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count. They retain the PBS key, not the repository password, so they can never open what the caller is asking about; serving them would have the agent try packages that cannot succeed and would let the screen claim an earlier package is openable on exactly the boxes the original defect hurt. The count is returned because their existence is load-bearing and underivable. The trade, stated rather than waved through: the hub still cannot read any of it - sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held. What widens is volume, bounded by self-scope, the recovery-mode gate and the cap. The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it; the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than 174. A positive control shows that check is name-presence, not decodability - filed as R-315 rather than reported as coverage. Six tests through the real endpoint; four red-proofs asserted applied. |
||
|
|
1d5f2b8bb6 |
DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption survived a week. (a) The material IS retained. host_escrow_superseded id 11 is the first retained row in fleet history to carry identity_blob (572 B), byte-identical to the pre-supersession row (sha256 a10032341c8584ed...). (b) The retained material DOES open the old store. Unsealed with the old recovery code it yielded a password byte-identical to the pre-change one, and restored three planted files byte-identical from a store the box itself could no longer open - including a Hungarian accented filename verified as raw bytes. Negative control ran first and failed closed. (c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero production callers; the recovery path selects FROM host_escrow. Asked with the code that had just worked by hand, the product answered "the recovery code did not open the sealed bundle". A valid code for retained history is reported as a bad code - the R-224 class again. R-304, rank 1. Both installer faults were watched happening first, so installer-v1.27.0 is now published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller 0.98.3 against a vouched 0.213.0, below the floor and below the version carrying the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own next install refused. R-297 and R-300 CLOSED. Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns on the second reinstall, proven), R-306 (--preflight-only writes state it says it does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 - operator decision), R-308 (stored controller password stale), R-309 (the day-0 runbook's publication claim has been false since R-110), R-310 (two edges). Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked operator-only. Phase A logs did not survive the intermediate revert; recorded. |
||
|
|
fbe1155fbb |
R-302 docs: register rows, the two rules earned twice, STATUS
gates / gates (push) Successful in 27s
Closes R-296 (verified: shipped in v0.212.0) and R-301 (premise confirmed, fixed in v0.213.0). Files R-302 with WHY the obvious condition was rejected, and R-303 for the missing markOrphaned guard - the co-render is now harmless, not impossible. Bake evidence for golden 0.213.0. |
||
|
|
890a474ff2 |
STATUS back to one screen; close R-280 and R-294
gates / gates (push) Successful in 21s
211 lines -> one screen. Moves closed items out, corrects the tester paragraph, states the floor situation as the operator's one-field call, and stops asking him to decide something that shipped. |
||
|
|
125aec1be2 |
R-300: uninstall no longer leaves dnsmasq blocking the next install
gates / gates (push) Successful in 18s
Removing the snippet and restarting left dnsmasq enabled and unconstrained on 0.0.0.0:53, so the next byo install's preflight refused and the customer went debugging a home network that was never at fault. Ownership is recorded at preflight (the only moment it is a fact - the package is installed by the agent, not this script) and honoured at removal. Boxes already in the field carry no record and fail safe to restart-only, with the reason and the command logged; the preflight message covers them instead. Not observed live - no installer-v1.27.0 tag is cut. Files R-299..R-301. |
||
|
|
f76cbf0ec7 |
Part 0: correct the PETI record - the mitigation it named does not exist
gates / gates (push) Successful in 20s
The row said a drive failure there means offsite-only recovery. Re-read from the hub's own store: no host row (deleted 2026-07-15 08:56:22, escrow_acked=0), no escrow of any kind, and offsite backup never ran once (escrow_state pending, snapshot_count 0 - the fork-4 guard working, not a fault). The local app-data repo was empty too and the whole-guest vzdump shares the failing device. If that drive fails today, everything on it is lost. Records the fact and leaves the parked/not-parked ruling open - that is the operator's call and does not need restating to be true. STATUS.md no longer reads the absence of a hub record as reassuring. |
||
|
|
eb600872f2 |
R-297: installer compares a local golden against the manifest before using it
Step 7 short-circuited on any local golden archive with no version compare, no digest and no warning, so the manifest sha256 was consulted only on the fetch path. Local discovery is newest-by-filename: correct by recency, never by verification. A box could reinstall from a stale archive and come back below the version where the offsite recovery screen exists. Digest first, then the baked controller tag. An auto-discovered mismatch re-fetches the vouched golden; an operator-named mismatch refuses. An unreadable manifest refuses rather than passing. Not published: installer-v1.26.0 is deliberately not cut until a fresh install has been observed taking a stale local golden on drill-r50. Also files R-295..R-298. |
||
|
|
11a5c3bd92 |
R-265 second sighting: a CI run failed and its log cannot be retrieved
gates / gates (push) Successful in 13s
felhom.eu run 293 ( |
||
|
|
c04f933d0b |
The census answers no, three receipts found, and the prune was on file all along
gates / gates (push) Successful in 13s
CENSUS (read-only, hub store, tester's machine not contacted): no machine that is not ours can be in the state that cost demo-felhom its history. The hub holds escrow for three hosts; both demo boxes lost their pre-fix key in the same four hours on 2026-08-04; peti-felhom and david have no host row and no escrow at all. A control ran FIRST and had to pass -- the query returned "present (572 bytes)" for a host known to have material and "absent (NULL)" for one known not to. Corrected my own instrument on the way: a date-only comparison mislabelled both losses as after the fix, so the in-force moment is now pinned from the hub's first post-fix escrow row (11:11:37Z), which independently agrees with the register. PART 1 ESTABLISHED. The prune is recorded inside R-267 -- the row about the Configuration page being slow -- because pruning artifacts is what made that page fast. Arithmetic checks (23+7=30, plus three versions that only surfaced after the first thirty moved them onto page one = 33) and the PAGINATED listing shows both generics at exactly ten. R-291's blocking condition is released: the operator was being asked to establish something already written down. And my counter-argument yesterday was wrong in exactly the way R-267 warns about -- "containers hold 19" came from an unpaginated query; paginated they hold 270 and 169. RECEIPTS: three restored (drives.enrol, backup.tier1, fail.lost-recovery-code), each citing the document that walked it; the map already read PROVEN-LIVE for all three, so this follows the map rather than raising a status in the view. NINE HONEST GREYS. fault.selfheal's best hit argues against it -- an incident recording self-heal's absence through a 1h15m outage. THE DECAY RULE FIRED FOR THE FIRST TIME. backup.restore-proof has a receipt from 28 July and is superseded anyway: demo-hp's restore-test failed 5 August and the box has since been rebuilt. A claim about a continuing behaviour cannot rest on an old observation. The capability map still reads PROVEN-LIVE and is now the thing out of step -- recorded, not silently rewritten. PART 4 specified, not implemented. The orphan card promises restorability the box rendering it cannot evaluate: the discriminator is on the hub and no wire field carries it. A conditional promise the system cannot evaluate is the same defect as an unconditional false one, so the copy stops promising, says what happens, and names a route. Ships with the next controller change so one bake covers both. |
||
|
|
67eced8fbf |
demo-felhom is protected again, and the authorised recovery could never have worked
gates / gates (push) Failing after 10m40s
Checked before acting, and the check is the finding. The box's local key and the hub's sealed escrow key hash to the SAME value (c60c8bc737a6b7c6...), and that key answers "wrong password or no key found" against its own repository. Running the recovery would have returned a key the box already held and which was already proven not to open the store. The store was written under 48741892f0ef4d59... -- host_escrow_superseded id=4, superseded 2026-08-04 07:20:08, identity_blob NULL. The restic password lives only in the identity bundle (escrow/identity.go:39, read by recover.go:91), so it is unrecoverable by construction; the surviving K-escrow payload is 64 bytes, a wrapped key, far too small to carry it. Same shape the register already records for demo-hp, four hours the wrong side of the retention fix. Took the operator's stated fallback instead: the orphan reset through the customer's own card. Old store moved aside, never deleted, to /home/felhom-repo.orphaned-20260810 (1.2 GB); fresh repository under the current key; offbox_repo_reset audited hub-side. Then PROVEN rather than assumed -- last_status ok, 10s, and the snapshot's CONTENTS listed: opengist compose files, manifest.json and volume-dumps/opengist_opengist_data.tar. Not an empty backup calling itself successful. R-202 gains hard evidence: the orphan card promises those set-aside backups may be restorable later with their recovery code. For these 1.2 GB that is false and unfixable, and it is said to the customers most likely to read it. |