e924a176323dbb1bed54ff7b1ecdcb19bbd67aa3
108 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cf09c78743 |
Slice 6 parts 0 and A: the guide in English, golden 0.258.0, and the CI finding
gates / gates (push) Successful in 30s
The volunteer guide has an English twin. It is a TRANSLATION, not a rewrite: 16 sections in the same order, identical step counts, table rows and warning blocks per section (measured, 0 sections differing in structure). Word counts are NOT a twin — English runs 19 % longer overall and up to 42 % on the short sections, because Hungarian is agglutinative; the +-15 % criterion the task asked for does not survive contact with this language pair, so structure is the measure reported instead. Golden 0.258.0 baked, published and vouched, with its record. One run, no aborted attempts: the 0.246.0 bake's two traps were both avoided by following its own record. Token proven not to leak with a planted control before the zero was believed. The waiver is NOT retired, and the record says why in one line: it is the mechanism of operator ruling R-468, not a note about this golden, and deleting it would turn the next release without a bake red immediately. It is also not load-bearing today. R-595: the catalog's copy gate could not run in CI at all — six pushes red, six alarm mails, while the local hook was green. Found by reading the operator's inbox, not by anything in the session that caused it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
31eeb36e88 |
ISO 1.29.0 PUBLISHED — bilingual console, proven on both menu entries (R-559)
gates / gates (push) Successful in 22s
Live at iso.felhom.eu, sha256 dceacae5da247d76cad065bf6c0d3bbefac8d8a5f8e571 db2fd8449a97e94829, and both download pages now name it. Every gate criterion is recorded with its OBSERVED value in documentation/tests/iso-release-1.29.0-2026-09-18/ — including two proof installs from the published bytes, one per boot-menu entry, each with a first boot AND one reboot: /etc/issue bilingual with zero hits for 8006, pvebanner masked, package 1.29.0 installed, unit enabled and fired, pairing code present, and the installed script byte-identical to repo HEAD. G11: the downloaded bytes hash to the published checksum. A defect was caught BETWEEN builds by looking at the screen rather than at the config: the second menu entry read "Felhom telepítés (szöveges mód) / Install Felhom (text mode)" — 58 characters — and the GRUB menu box cut it at "Instal". The English half was unreadable on the boot screen. Shortened to "… / text" and rebuilt; the published image is the rebuilt one. The Hungarian half is the part that may not change, so the English half is the part that gave. Teardown: VMs 323/324/325 destroyed, the two unclaimed appliance registrations discarded (zero left in `registered`), guest 9201 untouched — 23 containers before and after. The two stale *.rootpw.txt files were shredded from the publish source directory before the upload ran from it (R-587, files gone; the guard that would stop it recurring is still open). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
851198af7f |
golden 0.246.0 baked and vouched with agent 0.132.0; controller floor 0.246.0 (operator decisions)
gates / gates (push) Successful in 22s
Decision 1 delivered: floor 0.246.0 served, the N100 on 0.246.0 within seconds; Peti's box is DOWN on the hub and receives it when it reports. Decision 2: the agent vouch was refused by R-120 until a newer golden existed. On the operator's choice, golden 0.246.0 was baked (sha 05b7559d, amd64, all markers, token leak 0 with control 1, registry 200 before teardown) and vouched together with agent 0.132.0. golden_currency_gate: WAIVED -> OK. Bake evidence filed where the gate and runbook read it: tests/golden-0.246.0-2026-09-17/. Recorded, none reaching the registry: a first attempt on the arm64 template (my version sort), a self-matching pkill, and an OOM-killed watcher whose post-bake steps were done by hand. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
0a6cf60bf8 |
register: 9 closed (R-493/495/496/512/513/514/515/517/523), 9 opened (R-525..R-533), R-509/510/511/518 narrowed; ISO 1.27.1 publish record; website changelog; P1-fixes evidence
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
65790672d5 |
doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only, pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live (R-497 closed). Full first hour walked again on customer tester-1 (three disks + one disk): deploy, use, backup, removal, byte-identical restore, power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12 503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507, R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no longer claims the controller creates hostnames (R-506). NOT PUBLISHED. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
38848ffbeb |
drill: a stranger's first hour on 0.242.0 — 1 intervention, not ready for a volunteer
gates / gates (push) Successful in 21s
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468). Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps deployed and used, backup, remove, byte-identical restore, power cut and code typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494 (the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500 filed. Capability map: first-hour row added (PARTIAL), journey row scoped. Stopgap Hungarian volunteer guide written. Hub teardown layer pending. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ae59c31a84 |
R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day: - MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven` with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3 still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate restart logged "MariaDB upgrade not required" with the app serving. Evidence: documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine inside its major until Slice 4 (R-448) — removal tracked as R-469. - Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory, exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2, never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree. R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist. - Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0 (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was issued AFTER it landed. No --no-verify anywhere in this session. Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
417df06f35 |
slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing. |
||
|
|
bc47dd4ef9 |
v0.234.0: a known limitation written on 2026-09-02 was a defect by the next morning
gates / gates (push) Successful in 18s
The operator looked at demo-felhom and found OpenGist - up 15 hours, running exactly the catalog pin, showing no badge at all. 09-update-architecture.md had recorded that as an accepted limitation the day before: 'the fleet view fills in gradually'. On a quiet box gradually means never, and a feature that fills itself in on an event nobody triggers is, on the quiet installations, not shipped. That limitation row is now struck with the reason kept. The living document gains slice 1b, the two admission rules of the backfill (it never overwrites, and it refuses to seed a partial observation because the badge reads a service-count mismatch as BEHIND), and the note that the same field having two writers with two different admission rules is deliberate. Live evidence added: all nine apps already had records by the time 0.234.0 was ready, so the natural fleet state could no longer exercise the new code - said plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp by stripping two records; the backfill re-seeded exactly those two with digests matching independently-read ground truth and left the other seven alone. The refusal half was deliberately NOT staged live: it needs a degraded app, and manufacturing one risks the false-customer-email class that already cost 61 mails (R-330). Unit-tested with a red-proof, and recorded as unproven-live. R-457: a test that hardcodes a date and asserts an age derived from it is green only on the day it is written. Mine was, and it went red overnight. Six other files carry both a date literal and time.Now() - named as candidates, not accused. |
||
|
|
0705942783 |
the badge IS proven live, and the 'stale password' finding was mine, not the box's
gates / gates (push) Successful in 16s
I reported that the vaulted dashboard password no longer worked on either demo box, and quoted the controller's own 'Failed login' as the discriminator. The password was fine. ~/.config/credentials quotes its values with SINGLE quotes and my sed stripped only double quotes, so the quote characters went out as part of the password. The operator corrected it in one line; one retry returned 302. The instrumentation lesson is the finding and R-453 now carries it: 'Failed login' separates wrong-password from wrong-Host-header, and that is ALL it separates. It cannot tell a wrong password from wrong password HANDLING, and I read it as if it could. This is the second time this file's quoting has produced a confident wrong verdict, so the fix is one shared extraction helper, not a resolution to be careful. With the session recovered, the badge is validated on live pages: Naprakesz twice on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost (a deployed app with no record - absent is UNKNOWN, not current); and 'Frissites elerheto - 52 napja' on both surfaces, the age being real arithmetic on bentopdf's catalog_since. The behind state was staged by editing one compose tag, with no restart and no up -d, and reverted byte-identically (sha256 equal, diff empty, container never touched). Capability-map row upgraded to PROVEN-LIVE with the one unexercised badge state named. STATUS item 9 now needs nothing from the operator. |
||
|
|
6035dfcc3a |
09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the RestartStack comment that proves the restart half was CHOSEN (a design decision is not a defect), carries the three operator rulings of 2026-09-02, strikes the word 'rollback' (once a migration has run the old image will not start), states the target shape, and lists the seven slices with a status each. R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not changed. Nothing closed, so CLOSED-ITEMS.md is untouched. Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23 floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner), R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what stopped the badge render from being validated live). Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN LIVE through the boot reconciler on demo-hp - one entry per compose service, digests matching ground truth read independently. The badge RENDER is not, and the five attempts are listed rather than summarised. |
||
|
|
63eff21a5c |
golden 0.232.0 baked, vouched, floor raised — and it carries TWO releases
gates / gates (push) Successful in 17s
GOLDEN_SHA256 5f8a53ed5b19a6cb2006298ce6239f6fca2b990cc3ef6eada89f602801ca91b8, 657 494 489 B. 0.231.0 was never baked, so the fleet went 0.230.0 -> 0.232.0. THE CHECK THE 0.230.0 BAKE SKIPPED, AND THIS ONE DID NOT: the bake script's fingerprint was compared ACROSS THE HOP - 7b0fb5cf...73b6a1 on DooPlex and inside the VM. The previous bake recorded only the DooPlex-side hash and said so; this one is a measurement. Three independent readers agreed before anything was vouched: the bake's own print, the round trip of the published bytes (HTTP 200, 657494489 B, same sha, hashed from what was downloaded), and the hub's Day-0 dropdown reading Gitea on a different code path. The delivered artifact names its own controller - ./etc/felhom-controller-image reads felhom-controller:0.232.0 - with 19382 entries under var/lib/felhom/docker/. Both pre-gates were shown able to see something before their zeroes were believed, and the manifest was RE-READ after vouching rather than trusted from the 303 flash. DELIVERY WAS ACTUALLY EXERCISED. Both boxes had been hand-deployed during validation, so the floor had nothing to move. Rather than report delivery untested, demo-felhom was rolled back to 0.231.0 and the chain run for real - it moved itself in ~20s: 10:47:16 controller-swap: image file written, restarting bootstrap target=...0.232.0 10:47:26 controller-swap: new controller healthy target=...0.232.0 And this is the first bake golden_currency_gate.py actually gates: it now reads the GOLDEN_SHA256 line out of the bake log rather than matching a directory name (R-410, shipped hours earlier the same day). All 13 felhom.eu gates are green, golden-currency included, for the first time since v0.230.0 was released. |
||
|
|
7ee25925f9 |
R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp. LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level through the exact route the debug button invokes): - THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest declaring nothing - was pushed to the live store and the proof returned verdict "fail" with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes. - THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden: stopping the app does NOT produce a failed dump leg, because the off-site run's own capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no forget and no prune. State restored: the product's own run made a healthy snapshot the newest again and the proof then passed opengist. - The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s each, matching the spike's measured band. - The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear and vanish across a real restic check, and ZERO across the proof - including a direct 6x test of the snapshot-lookup argv, which settles that restic snapshots does not lock in 0.14.0 either. - Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the flag returned skipped:true duration_ms:0, no verdict, no alarm. - The customer's own verification copies were untouched throughout, which is the safety property the separate proof root exists for. ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at 19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the proof; I did not establish what it was. CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only - the job is REGISTERED, which is not the same claim. 07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore puts data back into a running app. Without that sentence the new green tick reads as covering the drill. REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 -> 152. No new rows minted. R-408 and R-409 stay open and are referenced by this work. golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job. A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses --no-verify for that reason - bypass #8. |
||
|
|
2263245cf2 |
golden 0.230.0 baked, vouched, floor raised - demo-felhom moved itself off the R-403 build (R-410 filed)
gates / gates (push) Successful in 17s
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e, 657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/. WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a good copy. Every fresh install and the whole fleet floor still carried it. golden_currency_gate.py had been red across |
||
|
|
83ff9e8e38 |
golden 0.229.0 baked, vouched, floor raised — R-242's sixth debt PAID the same day
gates / gates (push) Successful in 16s
GOLDEN_SHA256 39aa886df77b21757aef3b298a389343dc0df5134bb0f14e8f92a451d7bdae87, 656 864 331 B.
The evidence is the ROUND TRIP, not the build log: the published bytes were downloaded back and
match the bake on both size and sha, and ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.229.0 - the delivered artifact naming the controller it will start.
A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the
same sha straight from Gitea, a different code path.
Both pre-gates were proven able to see something before their negative results were believed - the
404 pre-gate against a 200 from 0.228.0, and the token-leak grep against a seeded throwaway copy.
Acceptance markers counted on the COMMITTED log: 1/1/1/1 present, 0/0/0 absent.
The vouch is a three-field change, checked rather than assumed: MinAgent 0.129.0 read from the
golden's controller CHANGELOG header, agent_version 0.130.0 >= min_agent 0.129.0 (not the R-216
shape), agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a
field it is not sent. Verified by re-reading the manifest, never by trusting the flash. The R-120
gate PASSED rather than being bypassed - fleet newest 0.229.0, golden 0.229.0.
The floor is proven ACTING, not merely set: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged
settle-gate GO at/above floor 0.229.0. Nobody deployed to that box. Both demo machines now carry the
Tier-2 unit restore.
golden_currency_gate.py went red -> green; the --no-verify bypass declared on
|
||
|
|
1623a4d5b5 |
golden 0.228.0 — baked, published, round-trip verified, vouched, floor raised
gates / gates (push) Successful in 19s
Bake: build-golden.sh v3.0.0 in the drill VM, reverted to virgin and cold-booted.
GOLDEN_SHA256 76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec,
658 079 744 B. All four acceptance markers counted 1; excluding/FATAL/mp1 counted 0.
The 404 pre-gate was proven to work before its 404 was believed — the target URL
404'd while the existing 0.227.1 package 200'd on the same command.
The evidence is the ROUND TRIP: the downloaded bytes match the bake's size and
sha, and ./etc/felhom-controller-image read OUT of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0.
Vouch: three fields together — golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from the controller CHANGELOG header, not assumed);
wrapper_sha256 carried through explicitly. agent >= min_agent, so not the R-216
shape. Verified by RE-READING the manifest, never the flash. R-120 gate passed.
Floor raised 0.227.1 -> 0.228.0 (impact preview {"below":3,"valid":true}).
The floor is ACTING: demo-felhom self-updated 0.227.1 -> 0.228.0 in under a
minute and re-registered offsite-integrity by itself. Both demo boxes now
re-read their whole off-site store on the weekly check.
Token hygiene: file->file scp, runner script inside the VM, unit properties
grepped 0. The leak grep on the committed log was proven with a planted copy
(1) before its 0 was believed. Teardown: guest 9100 purged, secrets shredded
after the log was copied out, VM off, disk reverted to virgin.
golden_currency_gate.py red -> green. All 12 felhom.eu gates OK.
|
||
|
|
db0812b6f2 |
Golden 0.227.1 baked, vouched, floor raised — and the floor delivered the new job by itself
gates / gates (push) Successful in 16s
Second full delivery of the day. golden_currency_gate.py went red -> green on the same command, so the --no-verify bypass declared on the previous push is now historical rather than standing. GOLDEN_VERSION 0.227.1 GOLDEN_SHA256 66754491dc9bd0130ef8ded9562f63c53a5ffdcfd91baa551141e55fa083ea32 size 657 403 203 B baked gitea.dooplex.hu/admin/felhom-controller:0.227.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP. The published bytes were downloaded back -- size and sha256 identical to what the bake reported -- and ./etc/felhom-controller- image was read OUT of the downloaded archive: felhom-controller:0.227.1. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Markers counted: docker OK (overlay2 = 1, mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. 404 pre-gate passed before the run and the script's own pre-delete agreed, so nothing was overwritten. Three-field vouch, all three checked: agent_version 0.130.0 >= min_agent 0.129.0 (NOT the R-216 shape), wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash. The R-120 gate on that POST passed on its own terms rather than being worked around. AND THE LINE WORTH KEEPING. demo-felhom self-updated 0.226.1 -> 0.227.1 in under 30 seconds and then logged: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST A box nobody deployed to now runs today's off-site integrity check on its own schedule. That is a floor DELIVERING rather than merely recording, observed instead of assumed -- and it is the strongest evidence R-242 has carried. Token hygiene: file->file, read inside the VM by a runner script, never on a command line (systemctl show ... grep -c -F token = 0). The leak grep on the committed log was PROVEN TO WORK before its 0 was believed. Teardown: guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out, VM powered off, disk reverted to virgin. R-242 now records the cadence as MEASURED: five convictions and two full bakes in one day. Every bypass declared, every debt paid -- and the pattern the row exists to name is exactly that a release and its delivery are separate acts. Its other half stays open: nothing gates the VOUCH itself. |
||
|
|
99af997ab9 |
R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163. |
||
|
|
4f875174fe |
Golden 0.226.1 baked, vouched, and the fleet floor raised — the debt is paid
gates / gates (push) Successful in 17s
golden_currency_gate.py had been CONVICTED four times today across three controller releases. One bake covers all three, and the gate went red -> green on the same command, which is its proof that it measures something real. THE THREE DECLARED BYPASSES ARE NOW HISTORICAL RATHER THAN STANDING. GOLDEN_VERSION 0.226.1 GOLDEN_SHA256 70ed8e9377dec22a9b493e55f222b0e25a49d7f3caec8c506e0412fd6baefe69 size 657 197 592 B baked gitea.dooplex.hu/admin/felhom-controller:0.226.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP, NOT THE BUILD LOG. The published bytes were downloaded back -- size and sha256 both identical to what the bake reported -- and ./etc/felhom-controller-image was read OUT of the downloaded archive: `felhom-controller:0.226.1`. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Acceptance markers counted, not eyeballed, each string captured from this run's own log rather than paraphrased from the runbook (two of the three the runbook named until R-233 could not match anything the script prints): docker OK (overlay2 = 1, including mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. The 404 pre-gate passed before the run, so nothing was overwritten. THE VOUCH IS A THREE-FIELD CHANGE AND ALL THREE WERE CHECKED: agent_version 0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash -- golden option 0.226.1 SELECTED, all four shas matching. THE FLOOR IS PROVEN ACTING, NOT MERELY SET. demo-felhom self-updated within 30 seconds: "[selfupdate] Post-update startup: update successful (0.225.0 -> 0.226.1)". Both demo machines now run 0.226.1 and only one of them was deployed to by hand. Token hygiene: copied file->file, read inside the VM by a runner script, never on a command line (systemctl show ... | grep -c -F token = 0). THE LEAK GREP ON THE COMMITTED LOG WAS PROVEN TO WORK BEFORE ITS 0 WAS BELIEVED -- a throwaway copy with the token appended grepped 1, was shredded, and only then was the real log's 0 taken as evidence. Teardown: build guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out (standing rule 5), VM powered off, disk reverted to virgin. R-242's OTHER half is untouched and still open: nothing gates the VOUCH itself. |
||
|
|
68a9f5475c |
hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s
One handler, two fields, opposite discipline. An unknown event_type is rejected with a loud 400. An unknown severity was rewritten to "info" without a word - and severityNotifies drops "info" before BOTH legs, so the event was stored, answered 200, and mailed to nobody. Two shipped features went out that way: DiskAlertKind.Severity emitted "warn" until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log rows before this session - not one, on any channel. The mechanism built to catch this class was structurally blind to it: the dispatcher's `unrecognized severity` line cannot execute for anything arriving over the API, because the coercion one line earlier guarantees the value it looks for cannot arrive. The coercion STAYS - a rejected event is a lost event, and losing an alarm is worse than mis-routing one. Only the silence is fixed: a WARN naming the customer, the event type and the rejected value. The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers, which never pass through the handler. For those it is the only severity guard there is. All 90 severity literals in internal/monitor are already valid, so the guard is silent because the producers are correct. Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the coercion test fails with "the hub rewrote a severity and said nothing". Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206. Vouching is the operator's act and was not done here. |
||
|
|
55274d5ef3 |
R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act. |
||
|
|
1eb64bec51 |
R-361 docs: the [FACT], the negative that cancelled Part 2, R-383/R-384, golden 0.221.1
gates / gates (push) Successful in 17s
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an invariant the code did not have, for four months - and a [DESIGN] on the db_dumps decision INCLUDING the trap it created: a stable list lets the already-current early return fire, so per-capture housekeeping must sit above it. 00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which IsDownState excludes. Measured on the shipped build with the scans demonstrably running over it. No suppression was built and no row opened. R-383: the double-failure message names an undo copy that is not there - R-361's own class, one surface over, observed on both 0.220.2 and 0.221.1. R-384: an app whose database has died reads unhealthy and raises no alarm. R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes. Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate blocked this push and that block is not circular, so it was satisfied rather than bypassed - no --no-verify anywhere in this session. |
||
|
|
a8caa0fdde |
R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step. |
||
|
|
8c9f1b798b |
golden 0.219.0 baked, published and round-trip verified (NOT vouched)
gates / gates (push) Successful in 18s
Baked in the drill VM per RUNBOOK-manual-build.md 4.0/4.1, carrying controller v0.219.0 (R-356). GOLDEN_VERSION 0.219.0 GOLDEN_SHA256 67b46f78f8ed9c7b1876265ab1bde9ec6798897898b1836acece9f3864a2aeb6 656832571 bytes All five pass markers matched, both negative controls at 0. Verified by ROUND TRIP - the published object downloaded again and its sha recomputed - not by the number the script printed. Both token-leak greps were proved able to convict before their zeros were believed: planted copy grepped 1, shredded, then the 0 accepted. Teardown complete: guest 9100 purged, four secret/script files shredded after the log was copied out, qemu exited, disk reverted to virgin. The revert first refused while qemu held the image, which is the runbook's own no-holder proof. NOT vouched - that is a three-field operator save (golden_version 0.219.0, agent_version 0.130.0, min_agent 0.129.0). |
||
|
|
877fcd2a38 |
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367. |
||
|
|
059adfb8b8 |
golden 0.217.0 baked and published — evidence, markers and the token-leak proof
gates / gates (push) Successful in 14s
GOLDEN_VERSION = 0.217.0 GOLDEN_SHA256 = 0276c5f638d140a861daba4ef25129896e259937eec0ae5cba7af42391315ad0 archive volid = local:backup/vzdump-lxc-9100-2026_08_21-21_40_05.tar.zst MinAgent = 0.129.0 (unchanged from 0.216.0) Acceptance markers counted in this run's own bake.log, not paraphrased: docker OK (overlay2 1 including mount point 2 (rootfs and mp0 — there is no mp1) upload OK (HTTP 201) 1 excluding 0 FATAL 0 Published package fetched back over HTTPS: HTTP 200. Token handling: copied file->file, read by a runner script inside the VM, never on a command line. systemctl show of the live unit contained it 0 times. The committed bake.log greps 0 for the literal token AND the grep was first PROVEN to work on that same file by appending the token to a throwaway copy (grep = 1) then shredding it — a 0 from an untested grep is not evidence. Evidence copied off the VM BEFORE teardown. Then destroy 9100 --purge, shred token+runner+script +log inside the VM (0 left), poweroff, waited for qemu using `ps -eo comm` (never `pgrep -f`, which self-matches), and reverted the drill VM to `virgin`. This unblocks the golden-currency gate, which correctly refused the previous push of the register rows: "controller v0.217.0 is released and NO golden carries it". No --no-verify was used. NOT DONE: the vouch. It is operator-gated and is a THREE-field change; vouching golden_version alone would ship this controller onto an agent older than it declares it needs. |
||
|
|
7d81681d6e |
golden 0.216.0: baked, published, vouched — gates green again
gates / gates (push) Successful in 13s
Closes the two-release day-0 gap that has been convicting CI since 2026-08-14. Run against RUNBOOK-manual-build.md 4.0 + 4.1. GOLDEN_VERSION = 0.216.0 GOLDEN_SHA256 = ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b archive = 656,970,239 bytes, controller image 0.216.0 template = debian-13-standard_13.6-1_amd64.tar.zst (listed live, not reused) Baselines re-read on the machine and all four matched the sheet: controller v0.216.0, its MinAgent 0.129.0, agent v0.129.0, previous golden 0.214.0. The published agent artifact for the vouched agent_version was confirmed present in the package registry rather than inferred from a CHANGELOG, and the R-216 check passed on the machine: MinAgent is EQUAL to, not above, the newest published agent. Verified beyond the script's own claim: the artifact was downloaded back out of Gitea and hashed, and it matches GOLDEN_SHA256 exactly. A script printing a digest and the registry serving those bytes are two different claims. Pass markers (corrected post-R-233 list) all present, quoted with line numbers in pass-markers.txt; excluding/FATAL absent; there is no mp1. Token never reached a command line: copied file->file, read inside the VM by the runner. systemctl show grep = 0. Token-leak grep on the COMMITTED log run with its positive control FIRST -- seeded copy 1, real log 0 -- because a grep -c that matches nothing also returns 0. Teardown: guest destroyed and purged, token/runner/script/log shredded AFTER the log was copied out, qemu exit confirmed with ps -eo comm (not pgrep -f), disk reverted to virgin. Vouched by the operator; verified by reading the hub's own store: golden 0.216.0 / agent 0.129.0 / min_agent 0.129.0, and the hub's recorded sha256 matches the independently downloaded artifact. That check was necessary because golden_currency_gate.py says of itself that it checks the BAKE, not the vouch. repo_gates.py --fast now rc=0, all nine gates OK -- first fully green run since 2026-08-14. Capability map deliberately NOT changed: the day-0 row cites drill documents, and the map's only golden literal is a dated historical citation on the recovery-journey row which bumping would falsify. R-334 is closed in a follow-up commit quoting this push's CI run id, since closing it without one would leave the ambiguity a third time. |
||
|
|
8b188bea68 |
hub v0.103.0 — a host can read the packages we kept for it (R-311)
gates / gates (push) Successful in 37s
ListSupersededEscrow had zero production callers for nineteen days. It is the only reader of a retained identity_blob, so the retention shipped in v0.93.0 was material the product could not reach - proven on the fixture 2026-08-12, where a code that opens a retained package was answered as a code that opened nothing. New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row GET, same recovery-mode gate, same audit event written BEFORE the bytes leave, capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count - they retain the PBS key, not the repository password, so they can never open what the caller is asking about, and serving them would let the screen claim an earlier package is openable on exactly the boxes the original defect hurt. The count is returned because their existence is load-bearing and underivable by the caller. The trade, stated rather than waved through: the hub still cannot read any of it - sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held. What widens is volume, bounded by self-scope, the recovery-mode gate and the cap. The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it; the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than 174. A positive control shows that check is name-presence, not decodability - filed as R-315 rather than reported as coverage. Also: golden 0.214.0 baked, published and round-trip verified; the countdown on demo-felhom cancelled on the operator's ruling (R-307); the spike that halted Part 3 recorded as R-312; the set-aside store found unrecoverable as R-313. Six hub tests through the real endpoint; four red-proofs asserted applied. |
||
|
|
fbe1155fbb |
R-302 docs: register rows, the two rules earned twice, STATUS
gates / gates (push) Successful in 27s
Closes R-296 (verified: shipped in v0.212.0) and R-301 (premise confirmed, fixed in v0.213.0). Files R-302 with WHY the obvious condition was rejected, and R-303 for the missing markOrphaned guard - the co-render is now harmless, not impossible. Bake evidence for golden 0.213.0. |
||
|
|
238954405c |
golden 0.212.0 baked and published; bake evidence
gates / gates (push) Successful in 23s
GOLDEN_SHA256=4b0a7dacc503c38732ed0a44949398639248c7fbd90758a1e4a047c21a7a15d8 Round-trip verified on the served bytes. Not vouched - the operator's. |
||
|
|
999b0a35f8 |
golden 0.211.0 baked and published; bake evidence
gates / gates (push) Successful in 22s
GOLDEN_SHA256=8593516889eb93fe1691410d7306be8cb87ee835b8d2378740eb34022272f849 Round-trip verified on the served bytes. Not vouched - that is the operator's. |
||
|
|
4a4a1e245a |
R-265 CI timeout + golden 0.210.0 baked; R-221/R-259/R-258 closed, R-266 minted, G-3 unblocked
gates / gates (push) Successful in 32s
Four defects of one family, all shipped today: something the box already knows, thrown away or drawn
as its opposite. Agent v0.128.0, controller v0.210.0. NO HUB CODE, no hub bump, no ArgoCD sync.
R-265 (this repo). timeout-minutes: 5 on the gates job — every honest run in the observed session
finished in 18-34s, so this is ~9x the slowest and far under whatever reaped run 264 at 834s with no
log. The alarm mail now carries Elapsed (start stamp via $GITHUB_ENV; an absent stamp prints
"unknown (no start stamp)", never a bogus 1.7-billion-second figure) and its "names itself in the run
log" sentence is qualified so it cannot mislead when there is no log.
⚠ THE UNKNOWN IS NOT CLOSED. Whether the if: failure() alarm fires for a REAPED job is still
unverified. The timeout makes the reap unreachable in practice; it does not answer what happens in
one. Demonstrating it means deliberately hanging a run on main, which would leave the branch red for
a parallel session. Said in the workflow comment, the changelog, R-265 and the report — none of them
claiming it is answered.
GOLDEN 0.210.0 baked, published, round-trip verified, NOT VOUCHED. The currency gate went red the
moment the controller was bumped — correct — and is closed by the bake, never --no-verify. No
--no-verify anywhere this session.
⚠ THE AGENT WAS NOT PUBLISHED UNTIL THIS SESSION CHECKED, AND IT MATTERED. R-221's fix is in the
AGENT, and a fresh install takes its agent from the Day-0 manifest. The binary had been hand-deployed
to felhom-pve and never published, so agent_version 0.128.0 was not selectable and a fresh install
would have received 0.127.0 — the golden would have carried the controller fixes and NOT the one the
headline defect needed. Caught by checking each Day-0 value was FETCHABLE rather than assuming.
Published from the live-deployed bytes, sha-verified across the hop first.
Registers. R-221, R-259, R-258, R-265 CLOSED. R-266 MINTED (READY): the failed root statfs still
travels to the hub as a 0-of-0 disk; ranked LOW because it is the quiet direction — it can only miss
a true alarm, never raise a false one — and it is now a two-repo wire change governed by G-1's gate.
Highest ID moved R-265 -> R-266.
CONTEXT S-39 rules the convention this project was missing: "we do not know" is never drawn as
"fine", and the codebase has ONE way of saying it — an explicit ...Known bool companion checked in
the template. ROADMAP G-3 was explicitly blocked on that decision and is unblocked; what remains
there is a survey-and-convert of existing sites, not the gate.
Capability map row 93 CHECKED and it was NOT claiming something untrue — it is about the operator
notification path. But its narrative ("the page you open to ask whether ONE app is backed up")
invites the wrong reading, and the adjacent thing WAS false until v0.210.0, so the row now records
that the two halves disagreed and only the operator half was true.
Six red-proofs across the two code repos, each with the mutation asserted applied. The one that
matters: Part 1 Scenario A FAILED against today's tree, with the intended message.
Part 1's operator-present live validation is OWED and is the session's STOP.
repo_gates --fast: all 8 OK.
|
||
|
|
dd55a3f98c |
correct the vouch state: the operator vouched 0.208.0 DURING this session
gates / gates (push) Successful in 33s
Recorded on arrival as 0.207.0, re-read from live hub_settings at the end and it is 0.208.0 — the operator acted while the session ran. The ask is therefore 0.208.0 -> 0.209.0, not 0.207.0 -> 0.209.0, and STATUS.md plus the golden evidence now say so. Caught only because the state was re-read rather than carried forward from the arrival note. A fact recorded at the start of a long session is a fact about the start of the session. |
||
|
|
3bf62b95bb |
fix(gate): the wire-contract search shelled out to grep and read its failure as a finding
gates / gates (push) Successful in 33s
CI convicted ALL 174 checked tags while the pre-push hook was green. Cause, read from the run log rather than guessed at the second attempt: the search used `grep -rnE --include=…`, and the CI runner's image carries python3 and git and deliberately little else — its grep does not support `--include`, so stdout was empty and the gate read empty as "the tag is absent". That is a gate silently treating a tool failure as a finding, which is worse than no gate, and it is exactly the error-swallowing this repo forbids. A green from it would have been just as untrustworthy as the red. Fixed by removing the dependency, not by working around it: the search is now pure Python — one token index per receiving repo, built in a single pass, no subprocess. Faster too (one walk instead of ~350 greps), and unreadable-file / empty-repo cases now exit 2 INCONCLUSIVE rather than reporting absence. THE BEFORE CAPTURE WAS RE-VERIFIED, NOT RE-GENERATED — the stronger claim. All 40 fields recorded in BEFORE.md were re-tested against the new implementation: agree=40, disagree=0, i.e. exactly the four this session fixed are now present and the other 36 still absent. The number 40 stands under both implementations; only the mechanism changed. The whole-token property survives by construction — a token index treats `healed_at` and `privsep_healed_at` as distinct tokens. This is the THIRD instrument defect this gate's own controls caught before it was trusted, after the substring false negative and the dr_recipe over-opacity. The first two were caught by re-finding the known instances; this one by the CI-versus-hook disagreement the workflow's alarm mail explicitly says outranks whatever the push was for. |
||
|
|
2ce3c2a0f2 |
golden 0.209.0 BAKED, PUBLISHED, ROUND-TRIP VERIFIED — the currency gate goes green (R-242)
gates / gates (push) Failing after 29s
The G-1 session released controller v0.209.0 (R-247), which made golden_currency_gate.py correctly red and REFUSED THE PUSH: no golden carried the newest release. The honest answer to that is the bake it asks for, not --no-verify. The gate's own docstring says the cost of a trip is one bake, which is the operation this project wants to be routine. 656 697 956 B, sha256 c9c4bcd6..e818ff. Round-tripped: the published bytes downloaded back, hashed independently, size and sha identical, and ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.209.0 — the delivered artifact naming the controller it will start. Acceptance markers all green (overlay2 x1, mount points x2 rootfs+mp0, upload HTTP 201 x1, excluding/FATAL/mp1 x0), Result=success, ExecMainStatus=0. 404 pre-gate with a 200 control on 0.208.0 so a 404 could not mean "wrong URL". Token file->file into a 0600 file read inside the VM; systemctl show grep = 0; committed-log grep = 0 WITH a control returning 1 to prove the grep works. Bake VM destroyed, /root residue clean, qemu confirmed gone, drill disk restored to virgin. IT ALSO CONSOLIDATES THE OPERATOR'S APPROVAL. Golden 0.208.0 was baked last night and never vouched; 0.209.0 contains everything it did plus R-247, so it supersedes rather than wastes it. One Save, not two — STATUS.md updated accordingly and back to its 93-line screen. NOT VOUCHED. Fresh installs still land on 0.207.0 until the operator saves. And R-242's untouched half showed itself again: this gate flipped green on the presence of the evidence DIRECTORY, with no vouch anywhere near it. Recorded, not built — ROADMAP G-8. repo_gates --fast: all 8 OK, including wire-contract and golden-currency. |
||
|
|
560f0d4451 |
G-1: a gate for the dropped field — built first, and seen failing on 40
Campaign 12 ranked this first of eight gating candidates. It is built BEFORE the fixes it finds, because last night an off-the-shelf tool for a neighbouring class (deadcode, for C6) was made to prove itself first and found NEITHER of the two defects it was meant for. A gate nobody has watched fail has not been shown to work. scripts/wire_contract_gate.py, registered in repo_gates.py as --fast (no network, no container, so it runs in BOTH the pre-push hook and CI — the R-29 constraint). THE TEST. For every json tag reachable from a declared wire ROOT, does that literal tag occur anywhere in the receiving repo's production Go or templates? A tag occurring nowhere cannot be decoded by any struct, named OR anonymous. That last clause is why a string test is used instead of comparing struct to struct: Campaign 12's first attempt paired types by shape and false-positived badly, because the hub decodes one report through several ad-hoc anonymous structs. RESULT ON TODAY'S TREE: 210 tags checked across 3 declared wires, 51 skipped (generic / opaque / allowlisted), 40 CONVICTED. Captured verbatim in documentation/tests/wire-contract-gate-2026-08-08/ BEFORE.md, which is deliverable 1 of this session. The prompt for this session said "465 emitted tags, eight unreachable". Checked against the repo rather than quoted: R-260's wording was "at least eight DECISION-BEARING facts", not eight tags in total. The real count on the three declared wires is 40, and R-260's own census already listed more than eight. Recorded because this prompt's own rule 6 says not to quote a document as source. TWO THINGS THE CONTROL CAUGHT, both before the gate was trusted: 1. A SUBSTRING FALSE NEGATIVE. `grep -F healed_at` also matches `privsep_healed_at`, so a genuinely dropped field read as received — and R-260 named healed_at, so its absence from the output was the tell. Now a whole-token regex; healed_at is convicted. 2. dr_recipe IS NOT WHOLLY OPAQUE. The hub stores each half as json.RawMessage and re-emits nested shapes verbatim, so the LEAVES are genuinely not on this wire. But the TOP-LEVEL SECTION KEYS are decoded by hostHalfShape/appHalfShape, and those are ALLOW-LISTS: a section an emitter adds is silently dropped until named in both. That already cost `offsite_restic` (R-122). So the gate is opaque BELOW depth 1, not opaque — the sections are checked and pass. Self-test: `--selftest` plants an unreachable tag on a real root in a throwaway copy and asserts conviction. Verified: exit 1, planted tag named. Blind spots are in the module docstring AND in the gate's own output, because Campaign 12's C1 guard turned out blind to one of the three shapes it was written for: generic tag names are not checked; reachability of a NAME is not use of a VALUE; only declared ROOTS are covered, and the hub's desired-state (served as raw stored JSON, no typed emitter) and the agent local API are NOT. Allowlist entries carry a stated reason. A quiet exclusion is a dropped field with paperwork. Not pushed alone: the fixes follow in the next commit so main is never red on this check. |
||
|
|
b7fb2117ae |
CAMPAIGN 12 — the class sweep: golden 0.208.0 baked (awaiting vouch), R-256..R-263 filed, gating ranked
gates / gates (push) Successful in 20s
Part 1. Golden 0.208.0 baked on the drill VM, published and ROUND-TRIP VERIFIED — 656 150 362 B, sha256 ba668f59..5ffb82, and ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.208.0. Acceptance markers all green (overlay2 x1, mount points x2 rootfs+mp0, upload HTTP 201 x1, excluding/FATAL/mp1 x0), Result=success. Token file->file, read inside the VM; systemctl show grep = 0; committed-log grep = 0 WITH a control proving the grep works. Bake VM destroyed, drill disk restored to virgin. NOT VOUCHED — the campaign halts there deliberately. golden_currency_gate.py was correctly RED on arrival and is green after the bake. No --no-verify was needed anywhere in this session. Parts 2-4. Seven defect classes swept for siblings by class rather than by feature. Analysis only: no product code, nothing deployed, no machine touched beyond the bake VM. Eight new rows R-256..R-263 (ceiling moved from R-255), grouped by class in OPEN-ITEMS.md. C1 produced no new instance and has no row. The sharpest is R-260: the agent reports operator_key_configured every heartbeat, the hub has no field for it, so the check that answers "can the operator get into this box" returns ok for a box with no operator key installed. Every class states whether its method re-found the known instances, because a method that cannot re-find them has not been shown to work: C1 2/3 (verified by replaying the pre-fix templates), C2 2/2, C3 2/3 + 1 as fixed, C4 fix-pattern re-found, C5 re-found, C6 deadcode 0/2 and bespoke 1/2, C7 weakest and said so. Blind spots stated per class; seven suspicions investigated and DISPROVED, including two of my own methods. Part 4's ranking is in ROADMAP.md as G-1..G-8. Gate C5 (cross-repo tag reachability — cheap, --fast-eligible, would have caught every R-260 instance on the introducing commit). Do NOT gate C6: golang.org/x/tools/cmd/deadcode was measured against a PLANTED probe and is blind to unreachable METHODS on widely-used types, which is exactly the shape both known instances have. R-242's untouched half is recorded, not built: this bake demonstrated it, the currency gate flipping green the moment the evidence DIRECTORY existed, before the round trip finished and with no vouch near it. Correction the campaign owed its own brief: escrow_stale was described as closed; it is R-247 and READY. The live repo is the source. Sampled rather than swept, exactly: C7 60 of 2652 production invariant comments and NONE of the 1440 test comments (that half is owed); C2 19 of 221 refusals; C3/C4 controller only. No finding was reproduced live. STATUS.md is 100 lines against its 93-line one screen. |
||
|
|
f651b31a7a |
golden 0.207.0 BAKED, PUBLISHED, ROUND-TRIP VERIFIED and VOUCHED — the currency gate goes green
gates / gates (push) Successful in 22s
Closes the delivery gap v0.207.0 opened this session. Until now the gate was correctly red and a machine installed today would have received 0.206.0 — the release written, tested and pushed, and not delivered. Round trip is the evidence, not the build log: the published bytes were downloaded back (656 879 192 B, sha256 20ec9602…22995, both identical to what the bake reported) and ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.207.0 — the delivered artifact naming the controller it will start. Acceptance markers were the ones R-233 re-captured from a real log: docker OK (overlay2…) x1, including mount point rootfs AND mp0 x2 (there is no mp1 since build-golden.sh v3.0.0), upload OK (HTTP 201) x1, excluding 0, FATAL 0. The 404 pre-gate ran WITH a control so a 404 could not mean 'wrong URL': 0.206.0 -> 200, 0.207.0 -> 404. The token never crossed a shell — copied file->file, read by a runner script inside the VM; systemctl show grep for the value returned 0. The token-leak grep on the COMMITTED log returned 0, and that 0 is evidence because a planted copy returned 1 before being shredded. Vouch was a three-field change with all three checked deliberately: MinAgent 0.127.0 read from the golden's controller CHANGELOG header, agent_version already >= it, min_agent not above agent_version (not the R-216 shape). Verified by re-reading the manifest rather than trusting the flash. The R-120 gate did not refuse. Drill VM restored to virgin; qemu confirmed exited with ps -eo comm, not a self-matching pgrep -f. R-242: the bake half is done and the --no-verify bypass declared earlier today is now historical. Its remaining half is UNCHANGED — nothing gates the VOUCH itself, so a baked-but-unvouched golden still passes the currency gate silently. |
||
|
|
c1dec41328 |
walk5 venue TORN DOWN — census 168 rows -> 67, and R-244 grew by 30 as predicted
gates / gates (push) Failing after 13s
Operator-confirmed. Stopped under a name guard (demo-hp carries its own 9201), aged past the hub's stale_threshold read from the DEPLOYED ConfigMap (30m), and polled delete-impact until deletable:true — treating an empty response as retry, never as success. Cascade + qm destroy --purge, guarded a second time. Every layer verified absent against a positive control that must survive and does: VM 300 drill-r50 and demo-hp's own guest 9201 still there; ep0 namespaces demo-felhom + demo-hp still there; wg peers .2 .3 .4 .250 still on the live wg0; hub rows for demo-felhom, demo-hp, peti-felhom untouched. 16.64 GiB returned against 17 G measured. RECORDED FOR THE NEXT TEARDOWN: the WG peer is removed on a ~5-minute SCHEDULE, not by the cascade. Immediately after the delete the hub row was gone while 10.77.0.5 was still on ep0's live wg0; wgsync had last run 37 seconds before the cascade, and the next push (4 peers) removed it, verified on the live interface at 16:57:07Z. The previous ledger checked this after it had already converged, so it read as instantaneous — a teardown that checks too soon would file a false finding. R-244 grew by 30 rows (app_log_issues), PREDICTED in the pre-run enumeration rather than discovered afterwards. Running total across torn-down venues ~101. Nothing here claims a clean teardown. Storage Box layer evidenced from the hub's own deprovision log: the HETZNER_API token in ~/.config/credentials cannot see box 611421 (subaccounts -> 404, storage_boxes -> 200 with 0 entries) — it is scoped to another project. |
||
|
|
3f4fb3825f |
R-201 CLOSED — the unaided recovery journey passes, both halves, on the fifth walk
gates / gates (push) Successful in 17s
Capability map: the unaided-recovery row turns FAILED -> PROVEN-LIVE, scoped, with what it still does not claim stated in the row itself: shape (c) did not fire positively (with the mint guard holding there is no local key, so the offer comes from shape (a)); and 'unaided' here means possible-without-a-shell, not obvious, because two obstacles are unsignposted. OPEN-ITEMS: R-201 closed with its evidence. Five new rows R-249..R-253 (the retrieval passphrase in page HTML; the host-key scan ladder vs AAAA settle; the listing's per-tag rows; the two unsignposted restore steps). R-243 annotated rather than re-filed: on a REBUILD offsite_delivery_stuck does not skip, so the row's gap is narrower than it reads. STATUS.md: headline changed, and trimmed 97 -> 92 lines rather than extended, per its own header. Teardown recorded as OWED with its before-measurements, the stop-and-age gate, and the positive controls that must survive. |
||
|
|
0691bc59a5 |
walk5 (R-201): Phase B + the verdict — BOTH HALVES PASS
THE DATA: PASS. All three sentinels byte-identical out of snapshot 5b0f20f7, including the accented filename's bytes, read back as bytes from the live path. THE JOURNEY: PASS — the first time in five walks. Zero guest command lines were needed to progress; the previous walk needed three. The reset-code hatch was used once, in Phase A only. §5's observation, which stands on its own whatever the verdict: at 14:58:52Z the rebuilt box collected its re-staged credential, configured the transport, and REFUSED TO MINT a repository password over the sealed package the hub holds. At the equivalent moment the previous walk minted a fresh key and lost the journey silently at 03:18. Sampled every 20s from T0: no key at any moment. Honest about which shape fired: with the mint guard holding there is no local key, so the offer comes from shape (a), not shape (c). Shape (c) was measured in Phase A in its NEGATIVE half (equal hashes, correctly silent). This proves the mint guard positively and the discriminator negatively. RTO 71.7s login to open store, of which 12.44s was the unseal and ~22s my own CSRF harness retry. Two new customer-facing obstacles, neither needing a shell but neither signposted: the restore refuses on unattached drives, and refuses because the app is not installed on a page that says the restore reinstalls it. |
||
|
|
879007aaec |
walk5 (R-201, fifth walk): Phase A journal — fixture built, sentinels proved by name
gates / gates (push) Successful in 21s
Written before the destruction, per §9.9. A fresh install landed on the VOUCHED set with no hand upgrade (controller 0.206.0, agent 0.127.0) — R-239's delivery gap is closed for this run, which is the first of the five walks where the box under test is the box a customer receives. Both §4.6 pre-destruction checks pass, neither previously exercised on a clean box: the recovery offer is correctly SILENT (shape (c) compares equal — the two key hashes are byte-identical on box and hub), and the restore page lists the app with the future-backup toggle OFF (R-237's fix, which the last walk measured failing). Also recorded: the §4.5 gate caught a harness fault (a toggle sent as enabled=1 rather than enabled=on) that had produced a green 'ok' over a zero-snapshot repository — the exact shape the gate exists for. |
||
|
|
721297ed5e |
golden 0.206.0 baked, published, round-trip verified — NOT vouched; the gate goes green
gates / gates (push) Failing after 10m43s
Controller v0.206.0 shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried 0.205.0, so a machine installed this morning would have received neither - and the next recovery walk would have measured the old behaviour and failed for a reason nothing to do with the walk. Same gap as R-239, one day after R-239 was closed. version 0.206.0 sha256 c85230b42f53baa9c1ee9986ac312c751d6cbc29fbe070d87bb2214429a9108e size 656,750,694 bytes (uncompressed 2,003,138,560) MinAgent 0.127.0 Round-trip verified rather than trusted: the published bytes were fetched back, re-hashed independently (match), zstd-tested, and ./etc/felhom-controller-image was read OUT of the download -> felhom-controller:0.206.0. That last step is the one that matters, because GOLDEN_VERSION is derived from the tag argument and could be right over stale content. All acceptance markers pass; unit Result=success ExecMainStatus=0. Secret hygiene: token file->file, in-VM runner so it never reached a command line (unit-property grep 0), literal-value leak grep on the COMMITTED log 0 - with a positive control proving the grep works before the 0 was believed. Bake VM torn down: CT 9100 purged, secrets shredded, qemu observed gone via ps -eo comm, drill.qcow2 reverted to virgin. THE GATE BUILT EARLIER THIS SESSION NOW PASSES. It was shown CONVICTED against the pre-bake state and is OK now - red to green on the same check, the same command, which is its proof that it measures something real. Note it went green on the BAKE, not the vouch: that is its stated limitation, and the vouch is still pending the operator. NOT VOUCHED - the operator's act. Only one field moves: golden_version 0.205.0 -> 0.206.0 (+ its derived sha). agent_version and min_agent both stay 0.127.0. wrapper_sha256 is unchanged but is CLEARED if omitted from the POST. |
||
|
|
094e93e828 |
finalwalk teardown complete; R-244 filed; session report
gates / gates (push) Successful in 13s
All five layers gone, each verified with a positive control that must
survive and does:
VM 324 + 4 disks -> absent (VM 300 drill-r50 remains)
hub: 13 tables at 0, incl. BOTH escrow tables (demo-felhom/demo-hp/peti remain)
Storage Box u629488-sub4 -> gone (sub1/2/3 remain)
ep0 PBS ns finalwalk -> gone (demo-felhom, demo-hp remain)
WireGuard 10.77.0.5 -> gone from the LIVE wg show on ep0, not just
the hub DB (.2/.3/.4/.250 remain)
14.06 GiB reclaimed against 15 G measured before deletion.
R shredded with a planted-copy control: plant -> search finds both ->
shred -> the same search finds 0. The zero was not believed until the
instrument was proven.
R-244 (NEW): a FULL census after the cascade logged COMPLETE full teardown
found 61 rows still matching finalwalk. Four sources are deliberate
provenance; the fifth, app_log_issues (29 rows), is NOT covered by the
residue purge - and it is systematic: c11 40, rewalk 20, part4 24 still
present from the 2026-08-06 teardown, whose ledger recorded zero
occurrences. That claim used a narrower query than a census and does not
hold; the correction is recorded in both the prior ledger and the register
rather than the measurement quietly redone.
No secret material is involved. The table is a fleet-wide aggregate: 12 of
the 29 rows are finalwalk-only orphans, 17 are shared with LIVE customers
and must be de-referenced, not deleted - very likely why the leg was never
written. Not fixed; a cascade change needs its own red-proof.
Lesson, and it is the reusable part: a per-table absence query is not a
census.
|
||
|
|
08b75e602e |
golden 0.205.0 baked, published and round-trip verified — NOT vouched (R-239)
gates / gates (push) Successful in 13s
Closes the delivery gap's build half. The vouched golden carried controller 0.203.0 while 0.205.0 was released, so a machine installed last night got neither R-237 (restore list keyed on the store) nor R-234 (skipped-app verdict). Both were measured from the customer's side on that box. version 0.205.0 sha256 8f49b2e8ccbc86a49df821fee9fb00c07293758811d3d0f0512dd0cf5fd54ee8 size 656,937,561 bytes (uncompressed 2,003,343,360) MinAgent 0.127.0 Round-trip verified rather than trusted: the published bytes were fetched back, re-hashed (match), zstd-tested, and ./etc/felhom-controller-image was read OUT of the downloaded archive -> felhom-controller:0.205.0. That last step is the one that matters, because GOLDEN_VERSION is derived from the tag argument and could have been right over stale content. Acceptance markers all pass; unit Result=success ExecMainStatus=0. Secret hygiene: token file->file, in-VM runner so it never reached a command line (unit-property grep 0), literal-value leak grep on the COMMITTED log 0 - with a positive control proving the grep works before the 0 was believed. Bake VM torn down: CT 9100 purged, secrets shredded, qemu observed gone via ps -eo comm, drill.qcow2 reverted to virgin. NOT VOUCHED - that is the operator's act. Only ONE field actually moves: golden_version 0.203.0 -> 0.205.0 (+ its derived sha). agent_version and min_agent both stay 0.127.0, because the new golden's MinAgent is also 0.127.0. The R-120 gate passes exactly: the newest controller the fleet reports is 0.205.0, so a 0.204.0 golden would have been refused. |
||
|
|
2228c0bff6 |
final walk COMPLETE — data PASS, journey FAIL; R-241 filed
gates / gates (push) Successful in 15s
THE DATA: PASS. All three sentinels byte-identical out of snapshot f5c53b03, including the 12 MB binary and the accented Hungarian filename whose NAME BYTES are identical too. Disk -> restic -> SFTP -> Storage Box -> rebuilt machine -> disk, intact. THE JOURNEY: FAIL, and further from the line than the previous walk. The claim worked first try (302 in 0.164s). Then: / lands on the launcher with no recovery pointer, /recovery 302s away, and the remote page offers to CREATE a new recovery code — which would orphan the history the customer's code protects. There is no field anywhere to enter the code they hold. The operator's documented remedy also refuses, correctly and fail-closed. Recovery needed three guest command lines. R-241 — and the cause is a success this same walk proved six hours earlier. OffsiteRecoveryOffer() shows the screen only when (a) there is NO repository password (pristine rebuild) or (b) one exists but the history will not open under it. Overnight the credential self-heal collected the staged credential and applied the tier, writing a FRESH key at 03:18Z — so (a) is false; and (b) is unreachable because orphan detection needs a run, and runs are blocked by escrow_state=pending. The gap is self-locking. Measured keys: on-disk 9b4a9a9d... vs recovered-from-R 30ef574f... This is R-218's shape one level up: succeeding at the self-heal stopped the box OFFERING the recovery it still needed. Registers: R-201 moved to its outcome; R-241 filed; capability map's recovery row stays FAIL with both halves and the cause named; STATUS rewritten for the operator. Highest ID R-238 -> R-241. The venue is left with the recovered key in place and the self-heal key moved aside, never deleted. Teardown still owed. |
||
|
|
1ff6f8f8e0 |
final walk §7: the credential chain runs end to end, unaided, on an UNCLAIMED box
gates / gates (push) Successful in 12s
Three questions answered from the hub's own log, not inferred:
1. the rebuilt, still-unclaimed box DOES report (host-report + Received report)
2. it DOES declare offsite.state=needs_credential, and offsite-delivery correctly
declines once a minute, naming internal/offsiteheal as the owner
3. offsiteheal re-staged UNAIDED at 03:15Z, after two reports carried the
declaration, with no provider credential minted — about 32 minutes after the
rebuild, matching the documented 2x15-minute debounce
And then the box COLLECTED it on its own 5-minute tick:
[offsite-apply] credential retry: the staged credential was collected and the
tier applied
That success line shipped in v0.203.0 and this is the FIRST time it has been seen
live: yesterday's walk only produced its sibling before I intervened at 102s and
mistook my own button press for the cause — the error that produced R-236 and
forced its withdrawal. Here nobody touched anything and the box was not even
claimed. R-218's consume half, R-236's withdrawal and the previous walk's dead
end 1 are all settled by one unattended observation.
|
||
|
|
1b490c8cbf |
final walk: destroyed, rebuilt, and HALTED at the claim screen
gates / gates (push) Successful in 27s
Destroyed 02:40:31Z (guarded on hostname — demo-hp also has a guest 9201), drives wiped to 20K with the mounts deliberately left in place because the surviving raw mount IS the R-220 condition. Reinstalled through the published day-0 path, installer v1.25.0 fetched live; Day-0 provision SUCCESS in 2m32s. R-239 measured a second time, from the other side: the rebuild landed on agent 0.127.0 (no downgrade, no hand upgrade — that half is right) and controller 0.203.0. The box a customer would recover on tonight also lacks R-234 and R-237. The machine is AT THE CLAIM SCREEN awaiting the operator. A claim code has already been requested through the customer-facing path and emailed, so the morning is paste-a-code rather than request-then-paste. The reset-code hatch was NOT used and will not be: it is a guest command line and would fail the rule the walk measures. Stated plainly in the journal: journey steps from the destruction onward are driven over HTTP from the appliance to the guest's island address, as a browser would; some instrumentation reads are guest command lines and are counted as such, but none changed state or was needed to progress the journey. |
||
|
|
502078bebf |
final walk: the soak DID catch a scheduled cycle; both directions recorded
gates / gates (push) Successful in 20s
Window 2 (01:56-02:36Z) caught the off-site nightly at 02:15:03Z — snaps 1->2, last_run advanced with no prompting. Window 1 had already caught the local legs (db-dump 00:30:25Z, tier + fill-watch 01:30:19Z). My 04:15 prediction was wrong in the other direction; recorded rather than quietly fixed. What should have fired and did not: NOTHING. Six registered jobs were never seen in the log and none is a finding — four are quiet by construction (scheduler.go:267, quiet := job.Interval <= 30s) and two run every 6h, outside a 4h window. Checked the source rather than filing four phantom defects. What fired and should not have: nothing. No alert, notification, email or digest. The only WARN lines in five hours were three of mine. Observation kept: offbox-backup ticked twice 62s apart and produced ONE new snapshot — the second silently dropped by the single-flight, which is correct for the nightly path and is the same mechanism that produced R-234 on the manual path. |
||
|
|
f873c557ca |
final walk: the soak, and the window correction stated as mine
gates / gates (push) Successful in 20s
Window 1 (22:44-01:54Z, untouched): every periodic job fired at its declared cadence — agent-channel-health 185x/1m, stack-scan 108x/2m, the 5-minute trio 44x each, hub-report 14x/15m. offsite-credential-retry ran 44 times and did no work and said nothing, which is R-218's asserted healthy-box behaviour on a box that never saw the defect. No alert, notification or digest fired; nothing on the must-not list fired. The off-site state never moved during that window — and that is NOT a finding. The daily jobs run on the controller's clock and the guest is UTC while the appliance is CEST, so the nightly local (~02:30) and off-site (~04:15) fall at 02:30Z and 04:15Z. I sized the window against CEST and it closed at 01:54Z, before either. Reporting 'the nightly did not fire' would have been a false finding produced by a badly-chosen window. Window 2 (corrective, 01:56-02:35Z) covers the 02:30Z local backup. The 04:15Z off-site nightly is deliberately NOT covered and is recorded as not run with the reason: leaving the machine at the claim screen before 07:00 is the primary deliverable, and waiting would have put the destruction at ~06:35 CEST with no margin. The off-site tier was exercised four times manually instead, including a full listing by name. |