1de6aaf904f7c811a0d2d93ba1c5ca6303daefd2
404 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
737694c603 |
CORRECTION: the app_oom alarm DID fire - R-635 was wrong, R-636 opened
gates / gates (push) Successful in 26s
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER skipped, correctly. I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape again, from the same hand, four days later. R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and emits once per container lifetime, so 4530 worker kills over six hours produced exactly one warning-level mail - indistinguishable from one transient kill. The magnitude was already collected (App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event. Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
8efd2d00df |
romm OOM storm on demo-hp: fixed, measured, closed (R-635)
gates / gates (push) Successful in 26s
Found because the operator heard the fans. romm 5.3.0 was promoted that morning; the update read `done` and the app ran clean for two hours, then OOM-crash-looped for six - 4530 worker SIGKILLs, ~500% CPU, host load 5.2 while otherwise idle, and nothing alarmed. Raising the limit to 768M was still a guess and fixed nothing (memory.peak hit exactly 768 MiB). Measured instead: ~216 MiB per warm uvicorn worker, so the image's default of 4 workers needs ~882 MiB. /init reads WEB_SERVER_CONCURRENCY; set to 2. Proven under load, not just at idle: 26,645 requests over 300 s, memory 416-614 MiB against 768, trending down, zero SIGKILLs, OOMKilled false. Idle CPU 500% -> 1.64%. The first soak measured nothing - it was pointed at the scratch-guest subdomain, every request 404'd at traefik in 9 ms, and the counter reported 14,026 successes. Positive and negative controls are now asserted before any load is driven. Carried into R-462: `proven` has meant "the update applied and the data survived", not "the new version runs". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
186546d562 |
THE TWENTY-EIGHT: every app no drill had touched, walked in one night
gates / gates (push) Successful in 27s
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a975cfde5b |
probe fix, the gate, and the promotion train (R-618 closed, R-630..632 opened)
gates / gates (push) Successful in 28s
Part 1: tandoor/zipline/wger probes corrected in the catalog and red-proofed live on 9202 in both directions - "Nem egeszseges" with the front door serving 200, then "Fut" after the real sync with no redeploy. tandoor's failed edge re-walked: done at +41.1s where it was failed at +361.9s. Part 2: fifteen proven versions on the live catalog, one commit per app; the guarded Update pressed on four apps on demo-hp, all four done. Opened: R-630 (paperless-ngx's probe has never run on any box - a silent absence, worse than the wrong probe that was found in one night), R-631 (five templates no static rule can judge), R-632 (28 of 53 templates never deployed by any drill). Closed: R-618. Register 318 -> 321. No product code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
462ab4a5ff |
Part 0: repair the register, gate its shape, and record Hetzner's answers (R-627, R-628, R-629)
gates / gates (push) Successful in 29s
THE REPAIR. Last night's append regex ate the state cells of R-446 and R-458, left them as a stray
fourth cell on duplicated copies of R-626 and R-625, and split the table with blank lines. The
register read 317 rows for 315 findings. Both cells restored from the cells that carried them, the
two duplicates deleted, and 15 blank lines that split the register into 12 separate markdown tables
removed. Every row's text is byte-identical afterwards, proven by diff; no row added or removed.
THE RED-PROOF FOUND AN OLDER INSTANCE: R-254 lost its state cell on 2026-08-08 (
|
||
|
|
8d786f7940 |
Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:` line is proven identical to before. WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN front door, with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured. THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, docker's own healthcheck green, and was then stopped and the household sent to a restore they did not need. zipline and wger are the same defect, both confirmed live. The gate that catches all three is static and cheap: both health checks already sit in the same file. WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) — which needed a purpose-built image store, because the rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. MariaDB across a major through the real button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted, with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at once and all ended honest. And the two EARLY power-cut phases nobody had cut in. TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English household in English and they do not, and R-446/R-458 are both narrower than their rows state. Two instrument fixes were needed before anything could be trusted: the unattended caller turned every success into a timeout (R-623), and one of my own reproductions was wrong and is kept labelled with what it actually measured. Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond the floor the operator asked for. Gates: repo_gates.py --fast, all 15 OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c85262111c |
The update arc's two missing measurements, the lock, and the floor to 0.260.0
gates / gates (push) Successful in 23s
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all down or blocked; both demo boxes SERVED. Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s after the cut decision and the seeded data read back intact — so the branch that is one step from old-binary-on-migrated-database is now evidence, not argument. Instrument limit stated: `starting` lasts under a second; all three landed in `verifying`, which RecoverUpdates handles in the same branch. Part 3 (R-611) — the night the previous session skipped without saying so. An app updated with nobody pressing anything; a terminally-refused app was pressed exactly once and never again over three passes. The unattended HOLD was NOT produced: the within-a-major rule correctly refused the broken edge before it was attempted, so Q4 still rests on the attended hold from slice 4. Said plainly rather than implied. Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard), R-614 (stale update phase survives a redeploy). R-520's pointer corrected. Catalog: two drill pairs, both reverted; every image line byte-identical to ff9717d3. The alpine:3.20 negative control a security review flagged is cleared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d19f07ea04 |
File R-607 (a sync that says 'no change' while the cache moves) and sharpen the audit
gates / gates (push) Successful in 26s
Found by the live run, not by reading: POST /api/sync answered 'nincs valtozas' while the box's catalog cache HAD moved, and catalog_images stayed stale until a separate rescan. Since CatalogImages is the one input CatalogOrder compares against, the badge answers from a stale catalog for that window — and the session nearly recorded a stale tag-ok badge as proof of the R-524 ahead arm. Neither half is isolated, so the row records the observation, not a diagnosis. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
0c263c77f2 |
Update arc resumed: the state measured, R-524/R-520/R-589/R-469 closed, seven questions put to the operator
gates / gates (push) Successful in 24s
Phase 0 — measured, never estimated: - both demo boxes: 10 apps, 0 behind, 0 unknown - 46 of 58 exact catalog pins are behind upstream; 39 within a major, 7 across - 6 of 7 measurable floating pins have been repushed since the catalog set them (R-446 is no longer theoretical) - the "23 of 66 floating pins" figure repeated in four places was STALE; recounted to 10, with the definition written down beside it Three claims in the brief corrected, named first: - R-589 was NOT open — it shipped in v0.258.0; only the row was stale - the chaos-night canary is NOT a defect — both gates refused to certify by design - the hub half of the report confirmed, with the nuance that the raw payload is stored whole, so Slice 7 is cheaper than the row implies Closed: R-524 (controller v0.260.0, proven live in both languages), R-520 (power cut during a REAL version change — the pin goes back, the app runs, the page says so), R-589, R-469 (MariaDB half). Filed: R-605, R-606. R-462's stale scope corrected. 09 gains §3 decision 10 (decided by CC unattended — operator may reverse), §3b with the seven questions in the decision shape, §6.2/6.3 the two open slices, and §6.4 an update night costed from R-462's real numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
bcdd5b2058 |
floor 0.259.0 raised; R-601 withdrawn as FALSE; R-604 filed
gates / gates (push) Successful in 28s
R-601 said demo-hp was unreachable. The operator looked at the hub and said it was online. It was, and had been up four and a half weeks, reporting every few minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan` at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub had carried the right address in every host report, and `ip neigh` on felhom-pve had .104 four lines above the .87 I quoted — I searched that output for the address I expected instead of reading it for the address that was there. Both ssh entries repointed and verified; nodes.md corrected, including that the tailnet route for this box does not exist. The hunt then found R-604, which is the real defect: demo-hp carried a per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. `managed floor SERVED` fires once per change by design, so a box behind a static override is silent for ever and its silence is indistinguishable from a box that already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes and its claim page now answers "Wrong or expired code" in English. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
e02bc03819 |
hub v0.119.0 — English households get English words for their codes (R-597); R-596/R-598 closed
gates / gates (push) Successful in 24s
The setup code and the owner passphrase now follow the household's language,
one word longer in English so the entropy never drops (setup 3 hu / 4 en,
passphrase 5 hu / 6 en). List and count are chosen together so a caller cannot
pair an English list with a Hungarian count. Hungarian is byte-unchanged.
Three claims in the row were wrong and are recorded as such:
- the RECOVERY CODE is minted by felhom-agent from the EFF list and has
always been English; the hub does not own it and no row was added.
- no claim mail states a word count; the only count wording was the bind
page's passphrase hint, whose English half is now count-free.
- the proposed phone-safe filter removes 68% of the list (5270 of 7772
words) and was measured, then declined, with the reason in source.
Also: guide_quote_gate binds the English volunteer guide's three quoted
messages to the controller's English bundle — nothing did, so the guide would
have gone on quoting Hungarian after the fix. Seven decoys, all convicting,
including the name-for-fact one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
732e9b9e0c |
Teardown verified on ep0, and the two periods I got wrong (R-600)
gates / gates (push) Successful in 25s
The customer delete cascade logged "full teardown" while the drill box's WireGuard peer 10.77.0.5 was still configured on ep0. Checked THERE rather than inferred from the hub, then watched until it went: gone about 6 minutes later. The mechanism is asynchronous, not broken; the log line claims a completeness it does not yet have. Both periods this session inferred from two log lines were wrong — the delete's staleness window and wgsync's push interval. A period read off two log lines is not a measurement, and both rows now carry what was actually observed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
91f047dfc0 |
Localisation slice 6: the guide in English, and a stranger's first hour (R-561)
gates / gates (push) Successful in 29s
Part 0 shipped as controller v0.258.0 (its own commit). Part A is the guide's English twin. Part B is the walk: a box installed from scratch that day, used by an English speaker following only the English guide and the screens. VERDICT: not yet ready for an English-speaking tester, because the claim page — the one screen between them and their box — is English chrome with Hungarian messages (R-596, P1). Everything else held: the download page, the bilingual console, all three customer mails, the bind page and its refusal, the dashboard's first language from customer.language alone, both app pages, the whole catalog, the language switch both ways — every one of them with zero Hungarian lines. One intervention (I1 = R-494, filed 2026-09-14); the stop rule was not reached. The walk exercised what 2026-09-14 could not: the graphical installer, the auto-reboot, and the mailed link and self-bind page end to end — that walk's H1 is closed, because this session had a mailbox. R-214 CLOSED as a side effect and seen rather than reasoned about: the console's last paint is now the bilingual "the box is linked" banner. R-516 does NOT close, and the item-by-item note says why: more than half its twelve items are about what a HUNGARIAN household reads, and an English walk cannot see them. It now waits on a Hungarian walk with a second drive. Rows opened: R-596 (P1, the claim page), R-597 (the setup code is three Hungarian words), R-598 (the Backup page's protection warnings), R-599 (a drill's teardown is blocked 30 minutes by report staleness and the 409 does not say so). Golden 0.258.0 baked, published, vouched, with its record. The waiver was NOT retired and the record says why in one line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
cf09c78743 |
Slice 6 parts 0 and A: the guide in English, golden 0.258.0, and the CI finding
gates / gates (push) Successful in 30s
The volunteer guide has an English twin. It is a TRANSLATION, not a rewrite: 16 sections in the same order, identical step counts, table rows and warning blocks per section (measured, 0 sections differing in structure). Word counts are NOT a twin — English runs 19 % longer overall and up to 42 % on the short sections, because Hungarian is agglutinative; the +-15 % criterion the task asked for does not survive contact with this language pair, so structure is the measure reported instead. Golden 0.258.0 baked, published and vouched, with its record. One run, no aborted attempts: the 0.246.0 bake's two traps were both avoided by following its own record. Token proven not to leak with a planted control before the zero was believed. The waiver is NOT retired, and the record says why in one line: it is the mechanism of operator ruling R-468, not a note about this golden, and deleting it would turn the next release without a bake red immediately. It is also not load-bearing today. R-595: the catalog's copy gate could not run in CI at all — six pushes red, six alarm mails, while the local hook was green. Found by reading the operator's inbox, not by anything in the session that caused it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
8e9401c6bf |
Localisation slice 5 CLOSED: the catalog speaks English and the floor is at 0.257.0
gates / gates (push) Successful in 27s
Part C shipped the same day the pilot was read: fifty apps in three pushes, 1 031 of 1 032 strings. The English Apps list shows ZERO Hungarian app descriptions across all 53 apps — the only Hungarian left on it is the "Naprakész" badge (R-589) and the language picker naming itself, which is correct. The Hungarian Apps list is identical to the pre-slice capture once the per-session CSRF token AND Docker's own "Up N hours" container string are normalised. Both normalisations are stated in the evidence rather than applied quietly — the second one moved because two hours of wall clock passed between captures, not because any copy changed. Fleet floor raised to 0.257.0 with the declared MinAgent 0.131.0, above the vouched golden so the declaration carries it. demo-felhom went 0.255.0 -> 0.257.0 by itself in under 12 seconds and THEN rendered the English tagline: the floor delivered the feature, not a version string. Rows: R-593 (papra describes a session-signing key as "the app's subdomain" — the one string left untranslated) and R-594 (the catalog gate can convict a retrieval promise but has no way to REGISTER a true one, which the shared vocabulary's design calls for). R-560 closed. 281 -> 287 rows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
f538a03bd5 |
Localisation slice 5: the catalog's copy model is MEASURED, not proposed (R-560)
gates / gates (push) Successful in 25s
10-localisation.md §7 goes from [DESIGN, proposed] to [FACT], with the numbers it was specced against corrected — and one correction chose an instrument rather than a footnote. The catalog has 1 032 copy strings, not 835. 832 carry a Hungarian letter, which was right. But the ASCII-ONLY Hungarian is ~120 strings, not three: „Aldomain" appears 53 times and „A szerver domain neve" 53 times, and the three the plan named („Igen"/„Nem"/„Nincs") do not occur in this catalog at all. An accent-only gate passes every one of them inside an English block — R-565's blind spot arriving again in a different repo. New §10.6 records what was proven live rather than reasoned about: the English pages show English; the seven Hungarian pages are byte-identical before and after the push apart from the per-session CSRF token; and a 0.255.0 box with the block synced onto it renders identically and logs no warning, in a 93-line window that contains the sync's own lines, so the absence is evidence and not a dead log. Rows R-589 (the update badge is Hungarian on an English page), R-590 (the data-folder card's backup promise, likewise, and it is a promise about the customer's files), R-591 (Stack.Copy() deep-copies five Meta fields and not the new I18n map — safe today, which is precisely why it is a row), R-592 (three defects inside the new catalog gate, closed the same session, each found by its own decoy). R-560 updated: Parts A and B done, Part C waiting on the operator's read of the pilot. STATUS asks for that read, and for the floor to 0.257.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
31eeb36e88 |
ISO 1.29.0 PUBLISHED — bilingual console, proven on both menu entries (R-559)
gates / gates (push) Successful in 22s
Live at iso.felhom.eu, sha256 dceacae5da247d76cad065bf6c0d3bbefac8d8a5f8e571 db2fd8449a97e94829, and both download pages now name it. Every gate criterion is recorded with its OBSERVED value in documentation/tests/iso-release-1.29.0-2026-09-18/ — including two proof installs from the published bytes, one per boot-menu entry, each with a first boot AND one reboot: /etc/issue bilingual with zero hits for 8006, pvebanner masked, package 1.29.0 installed, unit enabled and fired, pairing code present, and the installed script byte-identical to repo HEAD. G11: the downloaded bytes hash to the published checksum. A defect was caught BETWEEN builds by looking at the screen rather than at the config: the second menu entry read "Felhom telepítés (szöveges mód) / Install Felhom (text mode)" — 58 characters — and the GRUB menu box cut it at "Instal". The English half was unreadable on the boot screen. Shortened to "… / text" and rebuilt; the published image is the rebuilt one. The Hungarian half is the part that may not change, so the English half is the part that gave. Teardown: VMs 323/324/325 destroyed, the two unclaimed appliance registrations discarded (zero left in `registered`), guest 9201 untouched — 23 containers before and after. The two stale *.rootpw.txt files were shredded from the publish source directory before the upload ran from it (R-587, files gone; the guard that would stop it recurring is still open). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5a654ebc9a |
slice 4: bilingual console + the English download page is live (R-559)
The image is BUILT but NOT PUBLISHED, and that is deliberate: publishing to iso.felhom.eu is public and irreversible, and the runbook needs a proof install on BOTH menu entries plus a reboot against the uploaded bytes. That is supervised, so I stopped there. felhom-installer-1.29.0-pve9.2-1.iso, sha256 c67ceaa3…fb02, with every mechanically checkable criterion passing (G1, G2, G5, G6, G7, G9, G16 — including both payload files byte-identical to repo HEAD). The download pages still name 1.28.0, the image that IS published. Pointing them at a file that is not there would hand every reader a 404. A new site gate refuses the two pages naming different installer files or checksums, so whoever publishes 1.29.0 cannot update one and forget the other. Measured rather than read: the pairing banner is 24 rows on a 25-row console. One row of margin — so the height is now pinned, because two more lines push the HUNGARIAN code at row 5 off the top, and a banner whose code has scrolled away is furniture. R-587: two root-password files from July sit in the directory the public ISO is published from. Both 404 on the bucket (against a 200 control), so nothing leaked — but the only thing keeping them off is an --include pattern they miss by an accident of naming. A pattern that protects by coincidence is not a control. R-588: release records live in two different places, which made me wrongly conclude 1.28.0's gate had never been run. It had. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
eb1ae37095 |
ISO 1.29.0 source + an English download page (R-559 slice 4, source only)
gates / gates (push) Successful in 23s
THE IMAGE IS NOT BUILT AND NOT PUBLISHED BY THIS COMMIT. Publishing to iso.felhom.eu is public and irreversible and its runbook requires a proof install on BOTH menu entries plus the 16-criterion gate run against the exact uploaded bytes. That is the operator's step. The download pages therefore still name 1.28.0 - the image that is actually published - and a new site gate refuses the two pages naming different files or hashes. Three texts a person meets before any dashboard become bilingual: Hungarian block first, byte for byte as before, then English, inside the same frame. The pairing banner, the bound banner, /etc/issue (and the postinst's byte-coupled copy), plus an English half on the GRUB entries. The Hungarian is a GOLDEN, not a grep: test/golden/*.hu.txt were captured from the script at |
||
|
|
183727db9c |
docs: localisation slice 3 CLOSED — evidence, R-558/R-555 closed, R-585 filed
gates / gates (push) Successful in 25s
Part B's live proof is a LOG LINE rather than an e-mail, and the evidence page says why: severityNotifies drops `info`, so every safe converted producer mails nobody, and every converted producer above `info` describes something bad that is not true. Triggering one would mean a false record on a real box's timeline or a state change the fences forbid. The log line was built in v0.256.1 for exactly this, after finding there was nothing to look at on either side of the wire. 15:07:28 controller_started [hu-only] — Controller elindult (0.256.1) 15:08:25 controller_started [+household(en)] — Controller elindult (0.256.1) The Hungarian sentence is identical in both, and the hub stored the Hungarian in every case including the English-household push. R-558 CLOSED, R-555 closed with it. R-585 filed: six producers still send Hungarian only because their sentence arrives already finished from another package — `offbox_enlarge_blocked` matters most, since it has no hub entry so its raw sentence IS the household's whole mail. 10-localisation.md gains 10.4; STATUS rewritten for the operator with the two decisions left (raise the floor to 0.256.1; whether to rotate the demo password after R-584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
1637fa655d |
docs: slice 3 Part A live evidence, and two findings (R-583, R-584)
gates / gates (push) Successful in 22s
The live proof: the box's own "send test notification" button pressed twice, 74 seconds apart, on demo-hp. Reporting `en` it produced "[Felhom] Test notification / Dear Customer, ..."; switched to `hu` it produced "[Felhom] Teszt értesítés / Kedves Ügyfél! ...", byte-for-byte the v0.117.0 literal. The operator's copy is identical in both, which is the half worth stating. R-583 (closed, hub v0.118.1): the test mail was the one customer mail that did not follow the language, and it is the mail an operator would use to CHECK that the language works. The surface you would use to check a feature is the one most worth checking first. R-584 (open, P2): five probe scripts from slice 2's releases B/C/D were still in the guest's /tmp carrying the controller password INLINE. The rule to delete them exists, was loaded, and was not followed three times running - so the rule is not the mechanism. All shredded; whether to rotate the shared demo password is the operator's call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9167cf53af |
hub v0.118.0: the household's e-mails follow the household's language (R-558 Part A)
gates / gates (push) Successful in 23s
The hub has written every customer e-mail in Hungarian whatever the box was set to. The box has published its language since controller v0.247.0; nothing read it. Now it does. Nothing an operator reads changes. The Hungarian mails are byte-identical, and that is a diff rather than a reading: 56 goldens per language captured from v0.117.0 BEFORE any string moved, and all 56 Hungarian ones pass unchanged after every sentence was routed through the new bundle. - internal/i18n: flat bundle, 79 keys, hu authoritative + hu fallback, ceiling 0. - customerMessages/severityLabels are DERIVED from the bundle, so a sentence is written in one place and all 40+ tests that read those maps still work. - Language order: last reported -> created-with -> hu. reports.language defaults to EMPTY, never hu: "never told us" is not "chose Hungarian". - message_customer on POST /api/v1/event, additive and optional forever, for the sentences the box composes and the hub cannot translate. - The bind page is per-language, and its `expired` state stays Hungarian: it is the state an unknown token lands in, so rendering a real English customer's token in English would make the LANGUAGE answer what the TEXT refuses to. Two defects found inside the release: - R-581: the newest report was picked by received_at, which has SECOND granularity, so same-second reports tied and the winner was arbitrary. Ordered by the autoincrement id now. GetCustomers() still has the shape - row open. - R-582: the English copy-guard stems, ported word for word from Hungarian, convicted 141 honest sentences. The English claim is a phrase with a modal. R-555 closed: the language allowlist entry is out of wire_contract_gate.py. hub_copy_gate.py follows the sentences into the bundle - without that it would have scanned four files that no longer hold any customer text and reported success. Three new decoys incl. an innocent control. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
20aafc3dec |
ops: fleet floor raised to 0.255.0 — and it delivered on its own
gates / gates (push) Successful in 23s
POST /configuration/global-floor with min_controller_version=0.255.0 and the
declared min_agent=0.131.0 (R-472); 303 flash=floor_set, read back from the
form, not from the POST.
The proof is the N100: it was never hand-deployed and its own Docker reports
felhom-controller:0.255.0 healthy within five minutes of the save. Three boxes
remain below — all BLOCKED or DOWN, which is a floor being held, not a floor
failing; each takes it on its next check-in.
R-580 filed: curl's %{redirect_url} rebuilds the request URL WITH the --netrc
credentials in it, so the hub password was printed into the session's own
output. Nothing written to a file, nothing committed. The build-deploy skill
now carries the rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
4df2cd5174 |
docs: controller v0.255.0 — the globe fix on the sign-in-flow pages
gates / gates (push) Successful in 22s
R-579 filed and closed the same day: five shell templates loaded style.css with no cache-buster, so a browser holding the pre-0.254.0 file rendered the new globe unstyled; and the globe sat outside the card. - STATUS.md rewritten for the operator: what was seen, why, the third defect found while fixing it (version disclosure on the guest share page, caught by TestShareGuest_HeadersTilesNoAdminChrome), and the one decision left — raise the fleet floor to 0.255.0, with what happens either way. - 10-localisation.md §3: the shells' asset tag, and why the two guest pages get an opaque tag rather than the version. - Audit D: the parity diff (91 of 106 fixtures identical, every dashboard page among them) and the live endpoint evidence from demo-hp guest 9201. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
95c10954ae |
docs: localisation slice 2 CLOSED (controller v0.254.0) — R-577, R-578, and a probe rule
gates / gates (push) Successful in 23s
10-localisation.md §10.3: the saved notes follow the box language at write time, with the one-night consequence stated rather than hidden; the globe, and the table of WHO reads which page and where its globe posts — getting that wrong makes the button do nothing, which it did on /recovery until the live probe found it. Decision 6 superseded a second time; decision 8 (a claim carries the visitor's language) recorded. Decision 5 of §11's anonymous-surface line: changing what a VISITOR reads is within what an anonymous request may do; changing anything the household owns is not, and POST /lang can do only the first. R-578 — the deadlock, and why it is a row rather than a fixed bug: UpdateOffboxStatus holds the settings write lock while running its callback, boxLang() wants the read lock, sync.RWMutex is not reentrant. On a real box an off-site run would have hung FOREVER holding that lock. The symptom was a test suite going from 8 minutes to a 25-minute timeout. Fixed and guarded, but the guard covers one package and three helper names; the class needs a gate. R-577 — a guest share visitor still has no way to pick a language, and the household's setting is the wrong default for a stranger. Deliberately left, pinned by a test, and the operator's to decide because it is a promise the share feature makes. .claude/rules/live-probes.md, unconditional: never send a deploy request for an app that is not installed, not even expecting a refusal — the endpoint accepts first and validates later. Two sessions made that mistake in two days, the second WITH a prompt line forbidding it. A prompt is read once; a rule file is loaded every session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
08e9ddbd5c |
docs: localisation slice 2 release B (controller v0.253.0) — R-557 progress, R-575, R-576
gates / gates (push) Successful in 25s
All 179 Hungarian error messages carry a key; zero remain. 10-localisation.md gains §10.2: the
four properties util.MsgError had to have at once and the failure each one prevents, and the
plural rule as a BUNDLE rule rather than a per-call-site flag, with the answerable sentence and
what the other option would have cost.
Two instrument defects recorded rather than tidied away, because both shapes recur:
R-576 — the parity gate has a measured blind spot. The bulk converter dropped the continuation
of multi-line concatenations, damaging 7 producers, and the gate stayed GREEN: every surviving
fragment WAS a byte-equal base-commit literal, so its question ("is this text real?") was
answered yes while the CALL had lost half its sentence. Two behaviour tests caught it. The
general form: a structural gate over the TEXT cannot see a defect in the CALL.
And the script counting what was left was case-sensitive, so it said "0 remain" while five did —
R-565's shape inside the measurement. Every "no Hungarian left" claim in this slice is now made
case-insensitively and with both controls.
R-575 — the soft memory-overcommit warning has no error to carry a key and no language where it
is built, so it renders Hungarian on an English page. Named in the code, not hidden.
Live evidence includes a mistake I made and corrected: a probe of the deploy refusal INSTALLED
vaultwarden on demo-hp (the endpoint accepts before it validates), the same mistake the previous
session recorded. Removed through the product's own path with its data; verified gone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
ea4f6ab340 |
docs: localisation slice 2 release A (controller v0.252.0) — R-557 progress, R-566 closed, R-572..R-574
gates / gates (push) Successful in 21s
10-localisation.md gains §10.1: the measured numbers (1 120 base literals, not 1 141; 226 converted), the flash-as-key design and why a key must be resolved by the READER, word order through Go's explicit argument indexes rather than a second placeholder syntax, and the parity gate that makes "byte-identical" a measurement instead of a reading. Two claims the plan carried that live source disproved, both about the wire, both recorded: the country table is NOT on the wire (only codes are), and the hub does NOT always compose its own customer mail — it falls back to the controller's event message, which is why those 31 sentences stay Hungarian until R-558. That survey is handed to R-558 as its input list. R-566 CLOSED (four app-named page titles now carry a %s). New rows R-572 (two funcmap helpers with no English form), R-573 (the two channel-health banners arrive as finished Hungarian), R-574 (handler_debug.go mixes page copy with payload). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2bd6fbfac9 |
R-553 + R-563 CLOSED (controller v0.251.0): evidence, 10 §9 rewritten, R-569..R-571, slice-2 dependency
gates / gates (push) Successful in 24s
- audits/r553-2026-09-17/: live before/after on demo-hp (CSRF redacted), the 409 refusal, the hub health block, red-proofs for all five sites, the site-5 fixture diff, gates. - 10-localisation.md §9: the five decisions with what each reads now; the rule (a text signature may remain only where the text is not ours); the one legacy exception and its end date. - Register: R-553 and R-563 closed to CLOSED-ITEMS; R-569 (four API handlers match English words), R-570 (the legacy stale-note fallback + the slice-2 fence), R-571 (classifier and alert placement documented nowhere). R-557 carries the R-570 dependency. 263 -> 264 open. - STATUS, including the live probe that installed an app and was removed the same minute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5b6ade033f |
i18n slice 1 release C (controller v0.250.0): R-556 CLOSED, live evidence, R-565..R-568, decision 6 superseded
gates / gates (push) Successful in 23s
- audits/i18n-slice1-2026-09-17/C/: live before/after/en/back on demo-hp (CSRF redacted), hub report hu/en/hu, red-proofs, the switch fixture diff (89 of 89), green gate. - 10-localisation.md: §2.2 executeTemplateLang facts, §2.3 what stays Hungarian + the English test's ASCII blind spot, §3 decision 6 SUPERSEDED 2026-09-17, §5 formal ceiling 16 and its under-count, English retrieval stems; §10 slice 1 done; §11 decision 6 struck. - Register: R-556 closed to CLOSED-ITEMS; R-516 extended; R-565 (ASCII-only Hungarian invisible to the English page test), R-566 (three app-name page titles), R-567 (wizard nav highlight), R-568 (disk rows reorder). 260 -> 263 open. - Capability map row, STATUS (needs you: the floor delivers the switch). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
bc15153e0f |
i18n slice 1 release B (controller v0.249.0): live evidence, R-564, R-556 progress, STATUS
gates / gates (push) Successful in 23s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
85ded1f1d1 |
i18n slice 1 release A (controller v0.248.0): live evidence, R-563, R-556 progress, STATUS
gates / gates (push) Successful in 23s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
bc1a5db86c |
i18n: operator rulings 1b (banner + download page in scope, R-559 unblocked) and 7 (interface nouns translate)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5b7b2b22f1 |
i18n starter: inventory (script + audit), rows R-553..R-562, language allowlisted in wire-contract gate
gates / gates (push) Successful in 23s
Phase 0 of the localisation starter: i18n_inventory.py counts every customer-visible Hungarian string; the audit names six further claims in the prompt that live source disproved. Rows for the compare-not-show sites, wizard deletion, the wire-contract comment blind spot, and localisation slices 1-6. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2c96a84f12 |
chaos-night fixes: R-539 closed PROVEN-LIVE, the morning note, the report
gates / gates (push) Successful in 21s
R-539 closed: five real controller kills on demo-hp 9201 with the production 24 h window raised controller_slow_crashloop, and exactly one operator mail arrived (09:29:40Z). The fast brake never armed. Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed). unproven.py: 35 of 55 not walked, no number moved. The report names the brief's wrong claims first and one recommendation not followed: the controller floor was not raised - validated on one guest, a gap of my own found during validation, and a floor above the golden reaches Peti's box too. The operator's call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3c1882a0a4 |
chaos-night fixes: rulings recorded, guide reordered, three rows closed, two filed
gates / gates (push) Successful in 21s
Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and the two new event types with their audiences; 03 records the slow counter as built; 07 records the restore-record persistence as a REVERSED design for the restore record only; CONTEXT.md carries the day's rulings. Guide: the recovery code moves after the first apps and waits for the yellow bar. Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected). R-552 filed: an interrupted-restore notice for a removed app never clears - found in my own v0.246.0 after the release was built. Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a) live proof and its teardown. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
be99cf7c74 |
chaos night: R-550 corrected - I guessed four endpoints and all four were wrong
gates / gates (push) Successful in 22s
The row first claimed a restore leaves no record anywhere, citing four status endpoints that 404'd. All four were paths I guessed. The real route, read out of the restore page's own JavaScript, is /api/backup/restore-status and it exists. The corrected finding is narrower and better: the endpoint answers with the Go zero value (started_at 0001-01-01T00:00:00Z) and carries no 'last' field at all, while the page's own script renders '<operation> sikertelen.' from st.last.message. The restore record is in-memory only and does not survive the machine stopping - exactly the case a hard reset creates. The original wording is left visible in the audit with the correction beside it; the register row is corrected in place because a register must be accurate. The reusable lesson: I found the real routes by asking the controller for its own rendered links. Guessing produced four confident 404s that I then reported as a property of the product. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9f40dc3289 |
chaos night round 10: a restore leaves no record, and four of my instruments failed
gates / gates (push) Successful in 21s
The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.
R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.
Four instrument faults, all mine, all in the evidence:
* a 'nothing was logged' claim that was unfalsifiable when written - the log
stream holds zero lines before a reset;
* an on-disk check against /opt/felhom/data, a directory that does not exist;
* a household count reporting 0 lines and 0 failures when the truth was one
line and it WAS a failure - the runner now prints both operands;
* the disk guard was a TRANSIENT unit reporting 'active' all night, and was
absent from the reset onward. It is now file-backed and enabled, and its
script is copied off the box for the first time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
889310ec17 |
chaos night round 9: what a lost hub report actually costs, measured
gates / gates (push) Successful in 21s
With the injector corrected, the hub really was unreachable. The controller built its 23:08:42Z report, retried the push three times over 1m40.8s and gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued, which is correct: a report is a snapshot, not a fact. The box passed. 26 containers throughout, every front door serving, both the hub link and the host-agent link repaired unaided the moment the block lifted, no alarm fired and none should have. R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report cadence (15 min), so ONE failed push spends the entire budget. The measured gap was 29m59s - one second inside the alarm. A healthy, self-repaired box came that close to paging the operator. Also recorded: the injected cut is broader than its name - it severed the controller from its own host agent too, which a real ISP outage would not do. The caveat travels with rounds 7, 8 and 9. The event-drop path remains unmeasured, because no event was raised during any cut. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
e61aac1d8f |
CHAOS NIGHT: two enumerated gaps become rows in the same session
gates / gates (push) Successful in 22s
R-547 (P3): a disk that fills and empties between sweeps is never mentioned to anyone. The guest's root filesystem sat at 96% for ten minutes and no alarm of any kind fired - checked twice, once by the round's runner and once independently after the fill was released. disk_critical is defined at >=95% used, but the fill-watch is a DAILY sweep plus one check ~90s after a controller start, so a ten-minute window contains no check. The timing was almost comic: the controller restarted at 21:28 after the previous round's power cut, so its single opportunistic check ran about twenty seconds before the disk filled. This is the ladder working as designed, not a missed alarm - it is filed because the honest answer to "would the household be told?" is no, and that is written down nowhere. R-548 (P3): the whole-guest backup's LOCAL tier cannot fit on a small-system-disk box and retries on that tier for ever. A ~29GB source into a 14GB pve-root, measured falling at ~16MB/s - under four minutes to a full / on the nested PVE. The product's behaviour is correct throughout: it failed the tier, named it, scheduled a retry, its status surface agreed, and the off-site tier then succeeded from the same snapshot in ~8.5 minutes taking no local disk. What is filed is the loop: on a box this shape the local tier can never succeed. Honest caveat recorded in the row - the 32GB system disk is this drill's own fixture choice - but nothing checks whether the local target could hold the source before starting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
bca013edec |
CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
gates / gates (push) Successful in 23s
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being restored when the plug came out - returned healthy. The only alarm was controller_started, which is what the ladder expects for a 60-second outage: no node_stale (30 min threshold), no app_start_failed (90s boot grace). No false alarm, none missed. Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the loop died with the box and zero lines is not zero failures. The round also handed over immich's whole diagnosis. app_oom fired - "immich (immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the app and the exact container. That is why immich saw CONNECTION_CLOSED and crash-looped twelve times. It is added as tonight's line on the EXISTING OOM row rather than filed as a new one, because this project's standing finding is that those signals are invisible inside LXC guests and on this box the scan caught one. The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm feed, correctly labelled, the whole time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9fae6dfa98 |
CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
gates / gates (push) Successful in 23s
The schedule was drawn from seed 20260917 and written into the findings document BEFORE round 1, with its re-draw log. Phase 0 measured: - golden 0.245.0 baked, published (registry 200, not an exit code) and vouched; the box installed itself from the published ISO 1.28.0 and landed on it with no hand upgrade (controller 0.245.0, agent 0.131.0). - ZERO operator presses: the waiting self-bind mail worked, and the acknowledged -delete path re-issued off-site AND PBS-DR credentials by itself (pbsdr_auto_reissue) - the F-14 half nobody had watched happen live. - R-546 filed (P2): tonight's own guide sends the household to create the recovery code ~17 minutes before the box can do it. It self-heals; the bar urges them there the whole time. Measured on both sides, not inferred. - R-543 proven through its whole lifecycle on a fresh box: bar present while paused, gone for good once escrowed. - Known rows met and recorded, not re-filed: R-542, R-536's failure events. Also recorded honestly: three harness errors of mine (a script that announced "all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that turned eleven of my own 401s into what looked like product refusals, and a head -12 that hid a disk), and a near-miss where I almost filed a defect against a drive gate that was working and logging at DEBUG. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d124c77e17 |
R-543 closed: the household is asked for the recovery code (controller v0.245.0)
gates / gates (push) Successful in 21s
The tier-3 pause is the zero-knowledge escrow design and is untouched. What was missing was the ASK, while the backup page promised the copy that had never run. - VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and before the first app - what the code is, where, write it on PAPER, and that Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13. - day0-install.md A.2b: the operator step for a REBUILT box, which was missing. Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of "Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go). This is the correction to last night's "zero presses" note. - 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and 2 records that the household is asked from first login. - capability map: the first-hour row's last gap closed, with what it still does not claim (no volunteer has walked the ask from the written guide). - register: R-543 CLOSED with the live measurements; R-545 filed (nothing un-configures an off-site target). R-511 was already closed yesterday. - STATUS: the answered publish question removed (1.28.0 is live), readiness yes. - evidence: red-proofs, the two-box live validation, teardown on three layers, and both of my own mistakes in this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
832218dca4 |
ISO 1.28.0 PUBLISHED on the operator's yes; download page and R-535 updated
gates / gates (push) Successful in 20s
Uploaded with env-only credentials and verified by ROUND TRIP: the downloaded bytes checksum to a4cd9b6d…, identical to the built file, and the checksum file is served. 1.27.1 stays in the bucket; nothing was overwritten. The download page now names 1.28.0 with the published checksum (BOM preserved, site gates green). R-535 closes with an honest caveat: the new banner ships byte-identical to repo HEAD and the string is in the published payload, but it was never seen on a screen — the box bound itself while the walk was headless. Also corrected: the 1.27.1 heading still said NOT PUBLISHED although it went out on the big night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3f7ac8ee6e |
the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s
The capability map's journey row now carries the half it could never finish: five photos in, deleted the way a child would, the old route refusing and touching nothing, the off-site restore returning them, and them opening — sha256 identical, 5 of 5, with a negative control. Stated with it, because both are true: the bind needed ZERO operator presses (the box registered itself and used the mail the hub sent itself), but the PBS cascade needed ONE — the Re-issue press R-511 documents, which then succeeded because of this morning's ep0 grant. R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on day one — a fresh box waits at „Kulcsletétre vár" until the household creates its recovery code, and nothing asks them to, while the tier-1 row already promises that copy. R-544 records a log line that says „escrow deleted" where the effect is demotion to retained custody. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c18efc0610 |
teardown layer 1, a proper secret-leak check, and the fresh-box proof in both rows
gates / gates (push) Successful in 20s
VM 335 purged with its disks; demo-hp's own containers untouched; evidence pulled off the box before the destroy, with the one thing I could not collect stated (the agent journal — root SSH is refused on the appliance by design). The leak check redone properly: six real secret VALUES as needles against all 41 evidence files, planted positive control matched 6/6, committed evidence matched 0. The earlier „22" was the word „password" in labels — a word count, not a leak check. R-537 and R-538 now carry the fresh-box proof: the labels on a box where off-site is on, the refusal that pointed at the off-site route, and five photos returned byte-identical. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a73abf04db |
R-534 and R-511 CLOSED — the ep0 grant proven end to end on a fresh box
gates / gates (push) Successful in 20s
The rebuilt-customer case reproduced by itself: the WG-registration hook refused exactly as R-511 describes and named the Re-issue action. Pressing it then worked — reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two seconds later. No permission error. This morning the identical action returned „missing Datastore.Modify … status 255" and a 502. The only change in between is the narrow grant on ep0, and the narrowest role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only, and PBS has no custom roles. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
63e2de9b9e |
R-543: off-site ON by default is not off-site WORKING on day one
gates / gates (push) Successful in 21s
Measured on the fresh box: tier 3 sits at „Kulcsletétre vár" — the off-site copy is paused until the household performs the key-escrow ceremony, and nothing asks them to. POST /backup/offbox/run returns 302 and produces no snapshot; the controller log shows only offsite-credential-retry. That matters more after today, not less: the new default exists because a one-drive box otherwise keeps the household's files in no tier at all, and the tier-1 row now prints „Az alkalmazás fájljait a távoli másolat … védi". On day one that sentence promises a copy that does not exist yet. The good half, proven on the same page: tier 1 reads „DB + Konfig" with the new sentence, and „DB + Konfig + Adatok" appears zero times — R-537 holds here too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2687819195 |
the drive was fine and my instrument was not — corrected, and R-542 filed
gates / gates (push) Successful in 20s
I read „formatted but not mounted" off /api/disks/candidates and briefly held it as a product fault. The storage page — the surface a household opens — says the opposite and is right: Adatlemez, /mnt/felhom-drives/adatlemez, default, active, ext4. R-542 records the real (small) defect: that endpoint offers a REGISTERED, in-use drive under „initialize", with already_mounted null. The page filters it out, so no customer sees it; it fed a formatting flow and it misled a session, which is enough. Also recorded: off-site is LIVE on this fresh box by default — „Aktív — nincs kijelölt alkalmazás" — an hour after that default shipped, with nobody pressing anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
f2250ca31f |
register: R-537/R-538/R-536 closed, R-534's grant recorded, three rows opened
gates / gates (push) Successful in 20s
Closed with live proof on demo-hp (controller 0.244.0): the per-tier label and the restore refusal. R-536 closed with its red-proofs and the hub's two new event types. R-534 carries the measurement that matters: DatastorePowerUser is Backup+Prune only, PBS has no custom roles, so DatastoreAdmin at the datastore root for the hub's user is the narrowest grant that works. The row stays open until a re-issue is seen to succeed end to end. Opened: R-539 (a second, slower restart counter — the operator's ruling, for the nightly), R-540 (one pool box, no selection rule when it fills), R-541 (no path to move a customer between off-site boxes). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
0b63574293 |
drill 0.243.0: F9'' answers the supervisor budget question; machine torn down
gates / gates (push) Successful in 20s
Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of them accumulated, because the window is 15 minutes. Four restarts, zero pauses. So the brake catches a FAST crash loop and is blind to a SLOW one — a controller dying every 20 minutes is restarted forever, and the only trace is an info event that mails nobody. Measured, not changed: the options are written into R-531 for the operator to rule on. Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers 9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |