Commit Graph

90 Commits

Author SHA1 Message Date
admin eb1ae37095 ISO 1.29.0 source + an English download page (R-559 slice 4, source only)
gates / gates (push) Successful in 23s
THE IMAGE IS NOT BUILT AND NOT PUBLISHED BY THIS COMMIT. Publishing to
iso.felhom.eu is public and irreversible and its runbook requires a proof
install on BOTH menu entries plus the 16-criterion gate run against the exact
uploaded bytes. That is the operator's step. The download pages therefore still
name 1.28.0 - the image that is actually published - and a new site gate
refuses the two pages naming different files or hashes.

Three texts a person meets before any dashboard become bilingual: Hungarian
block first, byte for byte as before, then English, inside the same frame. The
pairing banner, the bound banner, /etc/issue (and the postinst's byte-coupled
copy), plus an English half on the GRUB entries.

The Hungarian is a GOLDEN, not a grep: test/golden/*.hu.txt were captured from
the script at 183727db9c before one English line existed, and the harness
asserts each banner's first N lines are exactly the golden. Red-proofed by one
changed byte, by an "a" planted in the English block, and by an over-wide line.

R-586, found on the way in: running the harness UNCHANGED at the base commit
failed two R-496 checks. The script paints with `>`, which truncates a FILE but
is a no-op on a console device; ISO 1.28.0's new bound banner (c033b3b) paints
straight after the pairing one and wiped it before the check read it. c033b3b
did not touch the harness, and nobody saw it because the harness is in no gate
and no CI run. Fixed with a FIFO; production code untouched. The harness being
ungated is still open.

The release gate's G16 required every Felhom string to be Hungarian and would
have STOPPED this publication. Operator ruling 1b of 2026-09-17 supersedes that
scope, so G16 is rewritten rather than waived: Hungarian FIRST, pinned by the
golden, each secret named once per language.

letoltes.html changes by four lines only. The English link is not in the nav -
the nav is a shared block site_gates.py pins across every page, and the gate
convicted the first attempt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 17:40:29 +02:00
admin 5b7b2b22f1 i18n starter: inventory (script + audit), rows R-553..R-562, language allowlisted in wire-contract gate
gates / gates (push) Successful in 23s
Phase 0 of the localisation starter: i18n_inventory.py counts every customer-visible
Hungarian string; the audit names six further claims in the prompt that live source
disproved. Rows for the compare-not-show sites, wizard deletion, the wire-contract comment
blind spot, and localisation slices 1-6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 14:50:29 +02:00
admin 832218dca4 ISO 1.28.0 PUBLISHED on the operator's yes; download page and R-535 updated
gates / gates (push) Successful in 20s
Uploaded with env-only credentials and verified by ROUND TRIP: the downloaded bytes
checksum to a4cd9b6d…, identical to the built file, and the checksum file is served.
1.27.1 stays in the bucket; nothing was overwritten.

The download page now names 1.28.0 with the published checksum (BOM preserved, site
gates green). R-535 closes with an honest caveat: the new banner ships byte-identical
to repo HEAD and the string is in the published payload, but it was never seen on a
screen — the box bound itself while the walk was headless.

Also corrected: the 1.27.1 heading still said NOT PUBLISHED although it went out on
the big night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 20:25:55 +02:00
admin c033b3b617 ISO 1.28.0 source: the console stops showing the pairing code once bound (R-535)
gates / gates (push) Successful in 20s
Measured 2026-09-16: 25 minutes after a successful bind AND claim the console still
showed the pairing code under a line promising the screen refreshes itself.

print_bound_banner is printed the moment the bind delivery lands. It does NOT name
the dashboard URL: the one-shot delivery carries the customer id, passphrase and
mode, not the domain, so naming an address would mean inventing one. The residue —
the console still does not reflect the later CLAIM, because this unit has exited by
then — is recorded in the changelog rather than implied away.

Not published: the built image needs the release gate and the operator's yes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:05:35 +02:00
admin 65790672d5 doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only,
pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live
(R-497 closed). Full first hour walked again on customer tester-1 (three
disks + one disk): deploy, use, backup, removal, byte-identical restore,
power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12
503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507,
R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no
longer claims the controller creates hostnames (R-506). NOT PUBLISHED.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 18:25:01 +02:00
admin 6fd8c87516 doorstep: console is Felhom's (ISO 1.27.0 source), passphrase hand-over copy (hub 0.113.0 source), rulings
gates / gates (push) Successful in 19s
Phase 0: the public ISO never auto-installs by construction (no answer.toml,
G1); the operator re-affirmed the interactive installer 2026-09-14.
- felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue
  (no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and
  paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first;
  fake hub now sends a pairing code (the banner was never tested, R-502).
- hub: created flash + Credentials block tell the operator to hand the phrase
  over; the self-bind mail names the operator (R-497). Tests red first.
- iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494
  narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned.
ISO_VERSION 1.27.0 (not built, not published).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 17:26:53 +02:00
admin 5e8a82c3c4 night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section
(R-471); target-selection.md names real paths (R-461); R-93 carries the
fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461
and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481,
R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note.
Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/.
2026-09-13 19:38:00 +02:00
admin 5ef0f52bcd Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).

Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:30:17 +02:00
admin ae59c31a84 R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day:
- MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven`
  with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3
  still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate
  restart logged "MariaDB upgrade not required" with the app serving. Evidence:
  documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine
  inside its major until Slice 4 (R-448) — removal tracked as R-469.
- Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver
  (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory,
  exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2,
  never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree.
  R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist.
- Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0
  (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was
  issued AFTER it landed. No --no-verify anywhere in this session.

Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains
decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 10:14:37 +02:00
admin 574f5df107 the decoy sweep: 29 gates read, 16 fooled, 10 fixed - and a gate that refuses the next one (R-421)
gates / gates (push) Failing after 17s
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five
instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone
looking at something else. The gates enforce every other rule in this project, including the rule
that findings must be written down rather than left in prose. Nothing had ever checked the gates.

METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a
directory with the right name and no bake log, a handler case that exists only in a comment, a note
whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached
by reading. Reading is how all five hid.

RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4
holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather
than called sound. A gate nobody tried to fool is UNKNOWN.

SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with
os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the
moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical
planted file, and are the control that proves the cause was the listing and not the decoy.

IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a
marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it
carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it
cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session
that closed them the day before, and every one was exempt from the only check that reads that file.
The rows were repaired first and the conviction added second - registering a failing gate refuses
every push.

THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption.
It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a
DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the
very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426).

NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate:
R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a
hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed
FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed
verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378
is the record of a machine getting it wrong.

FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves
nothing, and manufacturing a finding to fill a row is worse than an honest NO.

No product code. No version bump. No image. No golden owed. All four runners green.
Register: OPEN 172 -> 178, CLOSED 160 -> 161.
2026-09-01 12:39:45 +02:00
admin 1e6c387a0b R-404 CLOSED with the ruling; R-417 CLOSED by cause removal; R-418/419/420 filed
gates / gates (push) Successful in 17s
THE RULING WAS NEITHER OPTION AS FRAMED. Both offered answers - narrow the gate, or leave it and
write waivers - argued about the gate, and the gate was never the problem.

DIAGNOSIS, from live source: golden_currency_gate.py never looks at the push. It compares the
controller's newest CHANGELOG heading against this repo's bake evidence and returns the same
verdict whatever you are pushing - correct for a standing invariant, wrong as a push gate. And
controller_gates.py had NO golden-currency entry at all. So the repo where a release happens never
checked, and the repo that cannot create the debt was refused on every push. 18 of the last 24
pushes here touched no code - measured, and the new classifier agrees EXACTLY - most of them by
construction, because the controller's code is in one repo and its register lives in this one. SIX
of those 18 were bake records, so the push that PAYS the debt is itself documents-only: the gate
was blocking its own cure.

Not the waiver its docstring prescribes: that clause was written for a release nobody wants a
golden for. R-417 was a release we DID want a golden for, on a night the runbook forbade baking. A
waiver would have recorded a lie.

RULING: block the push that can create the debt, notify the push that cannot.

The gate's logic, exit codes and wording are BYTE-IDENTICAL. Only the consequence changed, for one
gate, on one kind of push, with a loud ADVISORY block so nothing goes quiet.

R-242's vouch half is amended in place to say it is UNTOUCHED and still open - a baked-but-unvouched
golden still passes both the gate and the new notice. Do not read R-404's closure as closing it.

FILED: R-418 - this runner's docstring listed ELEVEN gates while THIRTEEN were registered;
one-register and closed-register ran undocumented since 2026-08-24. Enumeration fixed here, the
correspondence is still unenforced. R-419 - observations_gate.py accepts an item whose body merely
CONTAINS "NOT-A-FINDING", even in prose disclaiming it; found by accident when a planted test
observation passed and my live validation proved nothing. R-420 - controller_gates.py could not
express a non-blocking gate at all before today.

Register: OPEN 171 -> 172, CLOSED 158 -> 160.
2026-09-01 12:01:10 +02:00
admin 6e550aedd3 R-87 put back in the register; closed_register_gate.py is the 12th gate (R-405, R-406)
gates / gates (push) Failing after 17s
Records and process only. No machine contacted. No product code, no version bump,
no build, no deploy.

R-87 was moved into CLOSED-ITEMS.md by the 2026-08-22 compression sweep ef6ac6f while
its own state cell read "READY - RE-RANKED UP 2026-08-03 (R-86 closed)". R-378 records
that sweep moving six still-open rows and restoring them in the same session; this was
a seventh it missed. Nine days in the wrong file, with the register's ranking paragraph
ranking it fourth and pointing at nothing. Restored verbatim from ef6ac6f^, beside R-95
where it sat before.

The predicate is the LEADING VERDICT of the state cell, which is R-378's lesson and
decides the answer here. Measured on the file as pushed: an open word anywhere in the
state cell convicts 3 of 151 rows, two of them genuinely closed (R-224 and R-260 carry
"open"/"OPEN" inside long prose verdicts); the leading verdict convicts exactly 1; the
whole row convicts 144.

closed_register_gate.py, two rules: no open state word leading a CLOSED-ITEMS.md row's
verdict, and no R- id with a row in both registers. Red-proofed both, and negative-
controlled against the pushed pre-fix files where it convicts R-87 by name, rc=1;
restoring the planted rows leaves the file byte-identical. Registered as the 12th gate
in repo_gates.py, --fast, after it was green. Four residual holes in its docstring.

R-398 was also in both registers - a deliberate cross-reference stub. Now prose beneath
the table rather than a table row, because a row in both files is what rule 2 convicts on.

R-406 filed: two unrelated findings in OPEN-ITEMS.md both numbered R-133. The only such
collision in either register. Deliberately NOT gated - a within-register duplicate rule
would fail on a pre-existing row, and a registered-but-failing gate refuses every push.

golden-currency is RED at this commit and was already red at dddcc80 - controller
v0.230.0 released, newest golden 0.229.0. Pre-existing, not this session's debt.

Ceiling R-404 -> R-406.
2026-08-31 15:32:40 +02:00
admin c30430c530 skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is
reported. Two rules this project has paid for — check the artifact rather
than the report, and do not state a claim more firmly than the evidence
allows — lived only in the operator's head and in chat, where Claude Code
never read them.

- felhom-evidence      five confidence tiers, artifact-over-report
- felhom-diagnosis     no hypothesis until a command has been seen red
- felhom-plain-language ASD-STE100, two options, the re-pitch
- felhom-handoff       the note goes to a FILE, not the conversation
- felhom-doc-authoring the pointer decides whether material is reached

scripts/check_skills.py asserts what decides whether a skill is EVER
reached: frontmatter parses, name == directory, description and body
non-empty, under 150 lines, installed copy still samefile()s into the
repo. install_skills.py globs and never reads the file, so a missing
description installs perfectly and then silently never loads.

It convicted on its first run: felhom-build-deploy is 179 lines. NOT
trimmed here (pre-existing skills are out of scope, and trimming a
deploy skill without exercising its commands is how a wrong command
reaches a live host) — a named single-entry GRANDFATHERED exception,
WARNed every run, R-394. A new skill over the limit is convicted.

Red-proof run and seen failing: description removed from
felhom-evidence -> exit 1, "frontmatter field 'description' is missing
or empty". Restored, tree clean.

skills/SOURCES.md records both MIT upstreams, that these are adaptations
not copies, and the six pieces deliberately EXCLUDED with reasons.

Register: R-392 (no architecture doc covers the two-AI workflow),
R-393 (decision-log skill deferred, with the reason), R-394.
2026-08-25 09:36:20 +02:00
admin ef6ac6fe74 One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)
gates / gates (push) Successful in 16s
Records and process only. No machine contacted.

ONE REGISTER (operator ruling). 17 roadmap rows moved into OPEN-ITEMS.md keeping their
identifiers, evidence and original filing dates - the oldest R-10, filed 2026-07-15, 38 days.
15 ideas stay in ROADMAP.md, which is their home; the gate exempts them by their own state
word. 59 already-closed rows stay as history. Sorting rule recorded in the roadmap header:
does the item assert something about the shipped product a reader could check and find false?

scripts/one_register_gate.py, wired as the 11th gate. Control run: baseline passes, a planted
open roadmap-only row is convicted by name, removing it passes with the file byte-identical,
and a planted `idea` row is correctly exempt. Its four residual holes are in its docstring.

The gate earned its keep immediately: it caught R-103, a READY finding my hand-sort mis-read as
done because my regex matched the whole row where the body contains "shipped" - the gate matches
the state cell. It also caught R-203 and R-163, recorded closed in the register and still open in
the roadmap; the roadmap copies are marked SUPERSEDED with the register's verdict.

HOUSEKEEPING. OPEN-ITEMS 672,376 -> 327,109 bytes (-51%); ROADMAP 239,306 -> 78,110 (-67%).
Closed work compressed to 17% into CLOSED-ITEMS.md and ROADMAP-HISTORY.md; every entry names the
commit whose git show returns the full original text. Rule-sentences are kept verbatim under
"Reasoning kept" rather than judged entry by entry - 25 carry one.

CONTEXT.md deliberately NOT compressed and the disagreement is argued in the report: 86% of it is
standing rulings still in force, this prompt's own 3.4 says the log is never edited, and it has no
per-ruling delimiter. Filed as R-377 - the problem is navigational, not volumetric.

The hot/bulk placement decision was NEVER recorded as a decision anywhere - established, not
assumed. Now marked [DESIGN] with a pointer honest about having no original date, given a
decision-log entry that records what was rejected, and the [DESIGN]/[FACT] legend carried from 1
of 8 architecture documents to 8 of 8. Existing statements deliberately left unmarked (R-376).

PROMPT-TEMPLATE gains N.7: compress what you closed, rehome live reasoning before it goes, state
the register's size before and after.

Ceiling R-375 -> R-378.
2026-08-22 12:13:54 +02:00
admin ca543b8f69 where-felhom-stands: bring the picture up to 2026-08-22, and stop the page disagreeing with its source
gates / gates (push) Successful in 15s
Eight claims re-checked against the drill and the v0.218.0 fixes; three moved, all downward.

  backup.offsite            walked -> partial. "18 snapshots, daily, unbroken" was true on
                            2026-08-09 and false by 2026-08-21: the next snapshot after that date
                            was put there by hand, twelve days later. The rebuild lost the target
                            and the per-app switches came back off, so a run reported "backup OK:
                            0 app(s) backed up".
  fail.wiped-reinstalled.data  walked -> partial. A real reinstall orphans BOTH off-premises tiers:
                            restic silently for 12 days (R-193), and the PBS archives from before
                            the reinstall cannot be opened by the rebuilt box at all (R-366).
  backup.fill-warning       walked -> partial. The warning fires correctly, but the watcher runs
                            once a day, so a filesystem that fills at 03:31 goes unannounced for
                            ~24 h. Watched silent while a volume sat at 99%.

Five re-checked and held: backup.tier1 and recover.byte-identical carry the R-355/R-354 story and
their fixes; backup.restore-proof stays grey for a sharper reason (orphaned archives, not an
untested tier); backup.sikeres gains two fresh instances; fail.customer-self-restore records that
R-356 now blocks 40 of 53 apps regardless of who is driving.

render_stands.py: the header's commit shas were hardcoded, so the page cited the August 9th commits
while the YAML said otherwise - the stale-build-product failure the renderer exists to prevent. They
are parsed now. The count beside them said "15 status(es) moved in that pass" when 15 was every
recorded move ever; it now separates the two numbers.

check_stands passes, and was itself proven able to convict first: a claim marked `missing` flipped to
`walked` in a scratch copy fired rule 5 by name (use.dlna), and the real file still passes.
2026-08-22 10:36:42 +02:00
admin 104ef34f57 gates: the today-override announces itself; malformed no longer swallowed
gates / gates (push) Successful in 15s
Part 6 of the hub-blindness task, separable and done rather than dropped.

Both due_checks_gate.py and instructions_gate.py read FELHOM_GATE_TODAY so
their suites can control "today", and neither said so. A shell that still has
it exported -- exactly what a session doing gate-test work leaves behind --
made both gates evaluate against a fabricated date and pass in SILENCE. That
is this project's own named failure class: an instrument that can quietly
return the wrong answer is not a measurement. The seam is legitimate and
stays; the silence was the defect.

Both now print a loud line naming the variable, its value, and that the real
date is being ignored, before any verdict.

And instructions_gate.py no longer swallows a MALFORMED override: it used to
fall through to the real date without a word while due_checks_gate.py already
exited 2 on the same input -- one variable, two gates, disagreeing about what
a mistake means. Both exit 2 now.

Tests extended in both suites (42 and 73 assertions, green). Red-proof: the
announcement was deleted from due_checks_gate.py and its two assertions were
seen failing, then reverted.

Also adds REPORT-hub-blindness.md (topic-suffixed; the shared REPORT.md is
left alone per the parallel-session rule).
2026-08-18 19:32:35 +02:00
admin 0a5e9b14dc due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
2026-08-18 15:16:51 +02:00
admin fc737b0fc0 installer v1.28.0: the removal genuinely reverses the installation (R-316)
gates / gates (push) Successful in 13s
v1.27.0's fix worked exactly once per machine. Measured on drill-r50 from virgin,
on the PUBLISHED v1.27.0, before anything was changed: cycle 1 recorded 'no' and
freed :53; cycle 2 recorded 'yes' and left dnsmasq running on 0.0.0.0:53; cycle 3
refused, exit 1. Every box already in the field is at cycle 2, and a reinstall
onto a machine that has had Felhom is cycle 2 by definition.

Why cycle 2 says yes: the preflight's ownership question is dpkg-query package
presence and nothing else - not the absence of a record. Stopping the unit and
leaving the package made our own package read as the household's one cycle later.

Now the uninstall removes the package when the record says we installed it. Order
unchanged and load-bearing: read the record, act, then delete the state file that
holds it. TWO packages are recorded, because dnsmasq ships the unit and
dnsmasq-base ships /usr/sbin/dnsmasq, and each is taken back only if we added it.

The dependency check is a SIMULATION, not a guess: apt-get -s purge is asked what
it would remove and the purge proceeds only if that set is a subset of ours;
otherwise stop+disable, naming the package that blocked it. Never interactive,
never fatal, and the success is re-queried rather than read off an exit code.

Watched: three fixed cycles -> install 3 PASSES; a household resolver untouched; a
dependent package not purged and named; no record -> untouched with the command
named. Red-proofs with the mutation asserted applied: remove the purge -> cycle 3
refuses in those exact words; remove the ownership check -> a household resolver is
purged; infer ownership -> the guess is taken.

Also: R-317 (the agent stats a path dnsmasq-base owns to decide whether to install
dnsmasq - pre-existing, now reachable), R-318 (no honest ownership marker exists
for existing boxes; the preflight message is the mechanism), and the status page's
decisions section rewritten to say what each decision costs and what doing nothing
selects.
2026-08-13 08:16:04 +02:00
admin 1d5f2b8bb6 DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption
survived a week.

(a) The material IS retained. host_escrow_superseded id 11 is the first retained
row in fleet history to carry identity_blob (572 B), byte-identical to the
pre-supersession row (sha256 a10032341c8584ed...).

(b) The retained material DOES open the old store. Unsealed with the old recovery
code it yielded a password byte-identical to the pre-change one, and restored
three planted files byte-identical from a store the box itself could no longer
open - including a Hungarian accented filename verified as raw bytes. Negative
control ran first and failed closed.

(c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero
production callers; the recovery path selects FROM host_escrow. Asked with the
code that had just worked by hand, the product answered "the recovery code did
not open the sealed bundle". A valid code for retained history is reported as a
bad code - the R-224 class again. R-304, rank 1.

Both installer faults were watched happening first, so installer-v1.27.0 is now
published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller
0.98.3 against a vouched 0.213.0, below the floor and below the version carrying
the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own
next install refused. R-297 and R-300 CLOSED.

Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns
on the second reinstall, proven), R-306 (--preflight-only writes state it says it
does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 -
operator decision), R-308 (stored controller password stale), R-309 (the day-0
runbook's publication claim has been false since R-110), R-310 (two edges).

Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked
operator-only. Phase A logs did not survive the intermediate revert; recorded.
2026-08-12 17:41:56 +02:00
admin 125aec1be2 R-300: uninstall no longer leaves dnsmasq blocking the next install
gates / gates (push) Successful in 18s
Removing the snippet and restarting left dnsmasq enabled and unconstrained on
0.0.0.0:53, so the next byo install's preflight refused and the customer went
debugging a home network that was never at fault.

Ownership is recorded at preflight (the only moment it is a fact - the package
is installed by the agent, not this script) and honoured at removal. Boxes
already in the field carry no record and fail safe to restart-only, with the
reason and the command logged; the preflight message covers them instead.

Not observed live - no installer-v1.27.0 tag is cut. Files R-299..R-301.
2026-08-12 14:03:40 +02:00
admin eb600872f2 R-297: installer compares a local golden against the manifest before using it
Step 7 short-circuited on any local golden archive with no version compare, no
digest and no warning, so the manifest sha256 was consulted only on the fetch
path. Local discovery is newest-by-filename: correct by recency, never by
verification. A box could reinstall from a stale archive and come back below the
version where the offsite recovery screen exists.

Digest first, then the baked controller tag. An auto-discovered mismatch
re-fetches the vouched golden; an operator-named mismatch refuses. An unreadable
manifest refuses rather than passing.

Not published: installer-v1.26.0 is deliberately not cut until a fresh install
has been observed taking a stale local golden on drill-r50.

Also files R-295..R-298.
2026-08-10 13:57:50 +02:00
admin 3ca9a7bbe6 R-242: build the gate the rule described — a release without a golden now fails the push
gates / gates (push) Failing after 13s
R-242 was filed 2026-08-07 as a mechanism-less rule and RECURRED WITHIN A DAY:
controller v0.206.0 shipped the R-241 fixes while the vouched golden still
carried 0.205.0, so a machine installed this morning would have received
neither. Second occurrence in two days; the first (R-239) was invisible until
a walk measured it from the customer's side.

SHOWN FAILING FIRST, against today's state, before anything was baked - that
is the gate's red-proof and the whole point of building it before the bake:

  newest released controller : 0.206.0
  newest golden baked        : 0.205.0
  GOLDEN CURRENCY GATE FAILED ... A machine installed right now would receive
  v0.205.0 - the release is written, tested and pushed, and NOT delivered.

Entry point exits 1; summary reports CONVICTED: golden-currency.

*** THIS PUSH USED --no-verify, to push past the gate's OWN conviction. ***
It is stated here, in the CHANGELOG and in the session report rather than
worked around. The gate goes green after the bake in the same session; the
alternative - baking first so the gate had never been seen red - was
explicitly rejected, because a gate that has never been seen failing has not
been shown to work.

IT IS --fast, AND THAT FORCED THE DESIGN. Both the pre-push hook and CI run
repo_gates.py --fast, which by contract selects only gates touching no
network. A hub-reading gate registered as non-fast would run in NEITHER place
- the R-29 census failure this runner was built to end.

SO IT CHECKS THE BAKE, NOT THE VOUCH. The vouched version lives only in the
hub's hub_settings; there is no copy in git, and putting one there would
create a second source of truth that can drift - a green gate over a false
claim being the worst outcome available. A bake without a vouch still passes.
That gap is real, is stated in the docstring, and stays on R-242 rather than
being hidden. The recurrence this gate exists for was a missing BAKE.

IT COMPARES VERSIONS, NOT BEHAVIOUR, so a release that changed nothing
customer-visible also trips it. Accepted deliberately: judging
"customer-visible" by hand is what failed twice, and the cost of a false trip
is one bake. A waiver belongs in the register, never in a habit of bypassing.

Inconclusive (exit 2) on an absent controller clone or an unparseable header:
not knowing is never a pass.
2026-08-07 12:53:43 +02:00
admin 15fa5273ba gate: check 7 (register citations) + content WARNings on the memory index
gates / gates (push) Successful in 8s
Check 7 catches "cites a register item and calls it open when it is not" -- the R-168 class, four
files, one self-contradicting. Trigger is an openness CLAIM, not any citation: policing every
mention would fire on ~30 legitimate provenance citations and the gate would be switched off.
Deliberate deviation from the task's literal wording, to keep it alive.

Two bugs found by the check's own red-proofs, both of which would have shipped:
- the state marker is not self-closing (**SHIPPED - text**), so the first parser read R-168 itself
  as OPEN -- a gate that cannot convict its founding case is decoration;
- the CLOSED exemption was line-wide, so "shipped" in a title pardoned "OPEN R-25b".

Check 6 gains WARN-only content classes on MEMORY.md. Link targets are stripped first: the earlier
scan reported three expired statements, all three false (dates in filenames), while missing the one
real expired claim, whose deadline was written ~08-02 with no ISO date.

39 -> 60 assertions. All four runners green.
2026-08-06 11:39:34 +02:00
admin f65ea89a24 workspace: version the root CLAUDE.md + InstructionsLoaded hook, and report which rules fire (R-229)
gates / gates (push) Successful in 8s
install_workspace.py lays down the two things that shaped every session while existing on one
host only. Unlike install_skills.py the targets are LIVE CONFIG, so: timestamped backup before
every write, settings.json MERGED (this script owns exactly one key), a diverged CLAUDE.md
reported rather than silently resolved, and an unparseable settings.json refused outright.

Proven: all 7 top-level settings keys survived byte-identically, and run 2 wrote nothing.

rules_report.py surfaces the column that matters -- rules that have NEVER fired, which are
mis-globbed or dead. 6 of 9 on first run. The hook now self-rotates at 5 MB.

The memory store is BACKED UP, NOT COMMITTED (auto-written, may name hosts/paths): added to
dooplex-backup.service's User Data component. /opt/backup/scripts/ is itself unversioned host
state -- filed, not fixed here.
2026-08-06 10:57:01 +02:00
admin f27aed87cd gate: instructions_gate check 6 — the auto-memory index (R-229)
gates / gates (push) Successful in 15s
MEMORY.md is the larger half of what loads before a word is typed (8.4k tokens vs the root
CLAUDE.md's 6.6k) and is the one instruction file nobody hand-edits, so nothing was watching it.

Three deliberately different outcomes, each pinned by a test: over-ceiling FAILS (auto-memory
drops content past the limit with no error), an orphan WARNS (the store is outside git), and an
absent store PASSES while PRINTING its reason -- asserted on the reason text, because a pass with
no reason is indistinguishable from a gate that stopped running.

39 assertions (was 20). Red-proof run against the real store, not a fixture.
2026-08-06 10:48:59 +02:00
admin 3a9dd81e18 docs+gate: felhom.eu/CLAUDE.md becomes core + path-scoped rules; instructions gate registered (R-229)
gates / gates (push) Successful in 8s
227 -> 115 effective lines, split into .claude/rules/{hub,website,manifests,docs}.md, and
repo_gates.py gains gate 6. Trim first, register second: a registered-but-failing gate refuses
every push through the pre-push hook, which is why this repo -- the one that OWNS the gate --
was the only one not running it.

Register discipline and the R-110 installer fence deliberately stayed in the core; both have
triggers no fixed glob covers, and scoping them would have rebuilt the failure class they exist
to prevent.

Scoping proven from the InstructionsLoaded hook log in two fresh sessions, not from frontmatter.
2026-08-06 10:41:42 +02:00
admin c21bcf84f7 docs+gate: instruction files cannot silently regrow (R-229)
gates / gates (push) Successful in 7s
New shared scripts/instructions_gate.py, registered in controller_gates.py and
agent_gates.py, never copied into a sibling repo (the reuse_refs_check.py
precedent). 20 fixture tests, all asserting the effect: exit code AND that the
message names the file and the reason.

It is a consistency gate, not a budget gate, and the failure message says so. A
/context reading measured the instruction files at 15k tokens against 869k free in
a 1M window -- space is not the constraint, and a future reader must not re-derive
the wrong reason. The 200-line ceiling is adherence guidance; a file nobody can
hold in their head is where contradictions hide, and five were found here.

Checks run against effective text (HTML comments stripped, because they are
stripped before injection): the line ceiling; every .claude/rules/*.md declares
paths: or an explicit unconditional: true; no component version literal; no
TEMPORARY block carrying a past date; and the workspace-root CLAUDE.md is
byte-identical to its versioned copy -- the live file sits outside any git repo,
so that copy is its only version-controlled record.

Two traps recorded so they are not reintroduced: a bare \d+\.\d+\.\d+ matches the
first three octets of every IPv4 (the gate excludes dotted quads, or it fails on
192.168.0.180 in the agent's own file); and unconditional: true is NOT a Claude
Code feature but this project's own marker.

Workspace-root CLAUDE.md 208 -> 182 lines (142 effective), copy kept identical.
The nine-instance invariant table moved into the felhom-testing skill, which
triggers when writing or reviewing a test; all three directive bullets stayed in
the core. felhom.eu/CLAUDE.md got surgical corrections only and is knowingly still
over the ceiling at 227 effective lines -- closing it needs the restructure R-229
defers, said plainly rather than quietly absorbed.

CONTEXT.md gains standing ruling S-35. OPEN-ITEMS.md gains R-229.

Docs only -- no Go, no version bump, nothing built or deployed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2
2026-08-06 09:38:52 +02:00
admin 51871a7ea6 installer 1.25.0: the off-site tier stops asking to prune (R-191)
gates / gates (push) Successful in 8s
Every weekly off-site run uploaded successfully and then failed the job on a prune
the box's token is deliberately refused — R-89 moved off-site pruning server-side
to ep0 and box tokens stay write-only. The 2026-07-26 'two weeks' ruling was not
reversed; where it is enforced moved, and keep_last: 2 did not follow. Now 0,
which the agent's existing guard already reads as 'never prune from the box'.

Verified read-only on ep0 before changing it: both namespaces have a prune job at
03:30 keep-last 2 that has run every day since 2026-07-27 — 18 tasks, all OK, the
newest keeping exactly two. Without that check this would have traded a weekly
false alarm for unbounded growth.

A gate asserts the offsite tier carries no client-side prune. The local tier is
untouched.
2026-08-04 09:46:09 +02:00
admin 688470c945 installer 1.24.0: a PRE-EXISTING backup target is granted too (R-185)
gates / gates (push) Successful in 7s
configure_backup_target has two arms and only one granted. Case A creates the
storage and grants in the same breath; the Scenario-F arm ('the target already
exists') returned without granting. A box whose felhom-backup pre-dated the
install therefore pointed local_backup_target at a storage its own token could
not read — measured on BOTH demo boxes: {"data":[]} through the token while root
lists three archives. That tier was never restore-tested and nothing said so,
because an empty listing is also what a brand-new tier returns.

The reuse arm now ensures the ACL through the same guarded wrapper. Scenario F is
unviolated: the storage DEFINITION is untouched, and pveum acl modify is
idempotent. BACKUP_TARGET_ID is deliberately NOT added to PVE_STORAGES — that
list is granted a step before the target is resolved, and --acl-storages entries
are preflight-checked for existence; the comment now says so.

A gate asserts it: every arm that resolves the target must also grant on it.
Red-proved by reverting the arm.
2026-08-03 18:58:49 +02:00
admin bee6848458 installer v1.23.0 — publishing becomes an act, not a side-effect (R-110, R-183)
gates / gates (push) Successful in 8s
Two channels moved off main in the same change, because either one left behind
makes the other cosmetic.

Channel 1 — the served script. webpage.yaml git-synced /scripts/ from
--branch=main every 30s and nginx served that tree, so pushing this file WAS
publishing it: within half a minute it was what every new machine downloaded and
ran as root, with no staging and no rollback but another push. The sync is now
SPLIT: the website keeps tracking main at the same cadence (a copy edit must
never need a release) and /scripts/ tracks the tag installer-v<SCRIPT_VERSION>.

PROVEN before the manifest was touched: git-sync v4.4.0 follows a tag AND
notices a MOVED one — measured on a throwaway sync against this repo,
"update required ... local:<old> remote:<new>" -> "updated successfully",
within one period. The moved-tag half is what the publish model rests on.

Channel 2 — the sixteen files fetched at run time. fetch_raw pulled from
$AGENT_REPO/raw/branch/main; it now pulls raw/tag/v$ART_AGENT_VER. That is a
correctness fix, not only a channel one (R-183): a fresh install fetched the
vouched agent BINARY while taking its unit file, sudoers and guarded wrappers
from whatever main held. Two refs, one install, nothing compared them. Their
correct ref was never SCRIPT_VERSION — they do not live in this repo.

No fallback to a branch: a vouched version whose tag is missing fails loudly
rather than quietly serving main.

Channel 3 — the URL — needed no change, recorded rather than left silent:
https://felhom.eu/scripts/felhom-host-install.sh never carried a ref, so both
producers follow the tag with no edit. No hub change, no hub version bump.

Gate 6 in hostinstall_gates.py pins all three structurally with no network, so
it stays in --fast and runs in CI. It deliberately does NOT assert "a tag exists
for the current SCRIPT_VERSION": that would go red on the very push that bumps
the version, before publishing — and publishing being separate is the ruling.
2026-08-03 12:08:37 +02:00
admin aa62449694 R-178 CLOSED: both demo boxes reinstalled from the merged golden and proven
gates / gates (push) Successful in 8s
Two boxes, two DIFFERENT supply paths, so the session proved the disk shape and
the delivery route rather than one of them twice.

demo-hp (layout proof, --golden <local volid>): mp0 at /var/lib/felhom,
backup=1, 70G, no mp1; /var/lib/docker and /mnt/sys_drive both real mounts of
its subdirectories via fstab; one df figure and one device id (64519) on all
three paths; reboots 3/3 with the binds surviving each.

demo-felhom (pipeline proof, --force-gitea-golden): fetch_verify succeeding
against the vouched manifest for BOTH artifacts -- 'verified sha256
54e2a4c431daf580... matches the hub manifest' for the golden, a7763d31... for
the agent. 250G single volume, grep -c '^mp1:' = 0, reboots 3/3.

Journey proven on both, endpoint-level: claim -> deploy -> back up -> restore,
with a planted marker returning byte-identical on each box. Ceiling measured
gone: 65 GiB and 233 GiB available to a recovery unit, against 19 and 45.

R-165 -> IMPLEMENTED, not PROVEN-LIVE, on the operator's ruling. B2, which that
row records as the bulkhead's replacement, fired live for the first time and
does refuse per app, delete nothing and alert -- but it is checked only in
captureAllRecoveryUnits while runVolumeDumps writes the bulk unguarded, and its
'the previous unit is untouched' claim was measured false (182,272 B dump
replaced by 2,147,666,432 B under a manifest still dated 06:34:26). -> R-181.

New: R-179 (uninstall leaves NAS network-storage units), R-180 (--archive-storage
not cross-checked against the ACL grant; 403 at step 8/8 after root@pam is
rotated), R-181. Third instance of R-115 recorded (agent 0.120.0 unpublished).

No code written, no version bumps -- this was a runbook.
2026-08-03 09:34:15 +02:00
admin e3525e62ac host-install: one data volume, derived from the disk (R-165)
gates / gates (push) Successful in 7s
felhom-agent v0.120.0 merges the two data volumes into one, and step_grows
computed two numbers while the install call passed both — so this had to
change with the agent or every install would have provisioned a half-sized
box. The 80/20 split is summed (226 = 184+42), so a standard appliance keeps
exactly the 250 G it had, no longer split by a wall.

The size still comes from the physical disk: step_grows already read the
thin pool's free space, and the merge only collapsed its two outputs into
one. --sysdata-grow is deprecated but still honoured, because the agent
folds a hand-passed value in rather than dropping it.
2026-08-03 06:43:58 +02:00
admin c718aad1bc docs: R-168 SHIPPED, R-29 CLOSED on the demonstrated alarm, R-169 minted
gates / gates (push) Successful in 7s
SPIKE-ci-runner-2026-08-02.md: all six probes with method, measurement and ruling; none
STOPped. P2 (stock image has git but no python3) and P6 (a runner that loses its state
re-registers and orphans the old record) changed the design; P5 (a failed run signals
NOTHING) is why the alarm exists at all.

R-168 SHIPPED with its evidence. R-29 CLOSED — on the demonstrated alarm and not on a green
run, as required: the class it opened is answered at both ends, the hook refusing locally and
CI catching a --no-verify bypass and emailing. R-161 noted: its automatic half now exists for
the STATIC gate, while its original scope, the runtime gate, is deliberately still not
automatic and should stay that way.

NEW R-169 (grep established R-168 was the highest in use): CI can only report, because there
is no gate in the road. Making it blocking needs branch protection plus a PR workflow, both
of which change how the operator works — so it is theirs to decide, and the row states the
cost honestly rather than recommending it.

CONTEXT gains S-8 (CI detects, does not block, and why that is structural), S-9 (a detector
that tells no one is not finished, plus the curl and Cloudflare-1010 traps), S-10 (the runner
is unprivileged because DooPlex is Tier 2), S-11 (CI reproduces the sibling layout).

CLAUDE.md gains the rule earned by red-proofing: a go test -run pattern that matches no test
prints ok and exits 0, and an instrument that can silently drop results is not a measurement.
2026-08-02 16:35:34 +02:00
admin 4707be755c docs: R-94 closed, R-29 leg (a) closed + leg (b) half, R-168 minted
hub/CHANGELOG v0.87.0 + scripts/CHANGELOG gate-enforcement entry. CONTEXT gains S-6 (the
hub renders no host-install version and the gate pins its absence) and S-7 (gates run from
one entry point per repo; reuse_refs_check was fixed rather than the REUSE.md convention,
with both rejected alternatives recorded).

OPEN-ITEMS: R-94 CLOSED all three legs, leg (a) by DELETION with its reason; R-29 leg (a)
CLOSED and leg (b) HALF-SHIPPED with the census result written into the row (13 gates; every
gate a CLAUDE.md names was green, two of the four unnamed were red); R-161 gains its
successor pointer. NEW R-168 (grep established R-167 was the highest in use): Gitea Actions
runner — measured 2026-08-02 as Gitea 1.26.2, Actions enabled on all four repos, 0 runners,
0 workflow runs, 0 branch protections, and the consequence that trunk-based direct-to-main
pushes leave no merge for a status check to gate, so CI here can detect but not block.
BLOCKED on a spike over host-mode vs privileged DinD on DooPlex and whether the workflow can
avoid JavaScript actions.

ROADMAP: R-94 collapsed to its one-liner, R-29 updated, R-168 added.
2026-08-02 15:28:31 +02:00
admin f2fc76ec4b ISO v1.26.1 PUBLISHED — both entries proven, round trip verified
Live: https://iso.felhom.eu/felhom-installer-1.26.1-pve9.2-1.iso
sha256 f3cc86d5f0ec68bba4155c994b4fa84e208d50209bb6e815636c99e5441059a6, 1705322496 bytes.

PART 5 PASSED ON BOTH MENU ENTRIES, four observables each:
  Graphical  spikegfx.felhom.eu  pairing code J7N-2DA
  TerminalUI spikesix.felhom.eu  pairing code ZY5-YY4
Both: manual install, own disk, own password, real completion signal, and the journal's
'not bound yet — polling every 30s ... normal waiting state, not an error'. Spike 4 had REASONED the
graphical path follows from shared Install.pm; it is now measured.

PART 6: G1-G10 + G13 all PASS against the uploaded file. G4's single hit is
felhom-bootstrap.sh:480's substring TEST ('$envtext' != *FELHOM_RETRIEVAL_PASSPHRASE=*), not a
value — my own regex matched the glob's asterisk.

PART 7: uploaded via rclone in a container configured ENTIRELY by environment variables, so no
credential file was ever written. Round trip verified from the public URL — not the local file.
Bucket stays private: unauthenticated GET to the S3 endpoint 400, custom domain has no index (404).

CORRECTED BEFORE UPLOAD: the generated manifest described a single automated entry with a 5s timeout
and listed Graphical/Terminal UI as 'menu-removed'. Generator fixed, sidecar regenerated, and the ISO
verified byte-identical before and after — the published file IS the file Part 5 validated.

Hub-side cleared: appliances 16, 17, 18 discarded (303 each); zero rows remain. The endpoint is
/appliances/<id>/discard, POST only (server.go:345) — not /delete.

Teardown: VMs purged, spike5 storage removed, demo-hp back to 6.6G, drill-r50 and 9201 untouched.

Still open and named: OPEN-ITEMS/ROADMAP dispositions for R-128/R-154/R-155 are not written; the
.deb is not byte-reproducible (G7 sub-clause); before-network stub unreached; Secure Boot and real
hardware not exercised.
2026-07-31 20:41:18 +02:00
admin 61e9b55737 SPIKE 4: a .deb in the ISO DOES deliver on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-4-2026-07-31.md

MEASURED, with a control, and the negative control is in the SAME box. One ISO (15 GRUB entries),
a trivial probe .deb injected into /proxmox/packages/, two qm-created VMs on demo-hp (400
interactive / 401 automated control) on a scratch dir storage at the /mnt/nvme-1tb mount ROOT.

Interactive (Terminal UI) install:
- package installed (ii felhom-spike4-probe 0.0.1)
- postinst RAN (marker + content intact)
- it enabled a systemd unit, and that unit FIRED ON FIRST BOOT (uptime 7.98s, pid1=systemd)
- while on the same machine proxmox-first-boot is NOT installed and /var/lib/proxmox-first-boot
  does not exist — Spike 3's negative reproduced, not assumed.

Postinst environment (identical both paths): pid1=unconfigured.sh, NO running systemd, but
'systemctl enable' SUCCEEDS; /proc+/sys mounted; network+DNS happened to be up (inherited from the
installer's DHCP — must NOT be relied on). Constraints: never systemctl start/daemon-reload, never
require network, never fail, do the real work in the unit at first boot.

Repack preserves it, but a naive 'xorriso -boot_image any replay' fails with 'Overlapping MBR
partition entries' — iso-repack.sh:270-292 already documents that exact failure and its fix.

R-153 RETRACTED into R-94 leg (b): OPEN-ITEMS.md:15 carries it verbatim at READY (XS), and R-29
says explicitly 'do not mint a new ID for a new instance'. Spike 3's further claim that the drift
leaves the generator 'three minor versions stale' was FALSE and is corrected — R-94 retracts that
exact reading; the served script is always main, so 1.22.0 is what every install already gets.

No new R-rows opened.
2026-07-31 15:35:07 +02:00
admin bb29186d62 SPIKE 3: [first-boot] does NOT fire on an interactive install
Findings only — no script, profile or build file changed; no release ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-3-2026-07-31.md

MEASURED with a control from the SAME image (one ISO, 15 GRUB entries):
- Automated entry  -> hook fires: ttyS0 marker, marker file,
  /var/lib/proxmox-first-boot/proxmox-first-boot (0700), activation symlink, unit active.
- Terminal-UI entry, normal manual install -> ALL absent, and the proxmox-first-boot PACKAGE is
  not installed at all. A whole-filesystem grep for the marker returns nothing.

Mechanism cited: Config.pm:118 defaults first_boot.enabled=0 and set_first_boot_opt is never
called in the Perl tree; Install.pm:746 returns early without it; Install.pm:1360 skips the
package. proxinstall (graphical) has ZERO occurrences of first-boot. [first-boot] is an
automated-installer feature, unavailable on every interactive path by construction. R-154.

A delivery mechanism DOES exist and is UNTESTED: Install.pm:1343-1372 unpacks every .deb in the
ISO's /proxmox/packages/ into the target on every path (fixed skip-list), then dpkg --configure -a
runs postinsts (:1378) — how PVE ships first-boot itself. Read from source, not measured.

Q5: the public image should carry NO answer.toml at all — that removes the baked root hash, the
disk profile and the whole Spike 1-2 problem space, and makes it a one-line release gate. But
iso-repack.sh:100-106 refuses an ISO without auto-installer-mode.toml. R-155.

Incidental R-153: hub hostInstallVersion=1.19.0 vs SCRIPT_VERSION=1.22.0; hostinstall_gates.py
detects it and exits 1 — the gate works, nothing runs it.

Q3 (real stub at before-network) was NOT reached and is recorded as not reached.
2026-07-31 14:47:18 +02:00
admin 19c932a693 SPIKE 2 complete: locked root closes the PVE web UI; before-network gives a measured zero window
Findings only — no script, profile or build file changed; no ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-2-2026-07-31.md

Both Tier 0 boxes went offline mid-session (provider cable fault; four routes tried, no Tier 2
fallback used) and returned. All three scenarios then ran to completion on real PVE, each signalled
by reboot-mode='power-off' rather than a disk hash.

- A LOCKED ROOT CLOSES THE PVE WEB INTERFACE. Measured at the exact endpoint the UI uses
  (POST /api2/json/access/ticket, root@pam) WITH A WORKING CONTROL: known-password install returns
  HTTP 200 + ticket; locked install returns 401 for every password and none can exist.
  passwd -S root = L, shadow = literal-asterisk, PVE uses the stock PAM stack.
- GRUB recovery mode is also closed ('the root account is locked') — but init=/bin/bash still gives
  an unauthenticated root@(none):/#. A locked box is recoverable, operator-only, at the console.
  The installed GRUB has NO password, so locking root is not a physical-security measure. R-152.
- before-network MEASURED (A/B, same image): the hook RUNS (marker, uptime 6.58s) with entropy 256,
  writable /etc, all binaries and openssl_rand_len=32, while ip_global is EMPTY and
  listen_22_8006 = 0. fully-up is the converse: sshd+pveproxy active, 3 listening. Zero window.
- R-148: answer.toml.tmpl:27 justifies fully-up with a pvesh/pct dependency the stub does not have
  (grep rc=1) — it blocked the ordering now measured as the fix.
- R-149 three ordering values; R-150 Condition-guarded hooks skip silently; R-151 demo-felhom built
  from an uncommitted profile.

Three probes failed and are recorded as failed: a container probe that ran as uid 0, a GRUB probe
that missed the 1-second menu timeout, and a kernel-line edit one line off (caught by a pre-typing
verification screendump). The interim 'Layer 1 teardown INCOMPLETE' is corrected — the fixture had
never landed, because the staging mkdir was in the SSH call that timed out.
2026-07-31 13:57:21 +02:00
admin 5bdd8372f8 SPIKE 2: before-network gives a zero window by construction; locked root closes sulogin
Findings only — no script, profile or build file changed; no ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-2-2026-07-31.md

BOTH Tier 0 boxes went offline mid-session (remote site, 12:28 CEST; four routes tried, our
tailscale pod healthy). Q1/Q2/Q3 each keep a part needing a nested VM: those are BLOCKED, not
answered. DooPlex was NOT used as a fallback — Tier 2, and this task did not authorise it.

Established without them:

- STRUCTURAL: ordering='before-network' maps to proxmox-first-boot-network-pre.service
  (Before=network-pre.target, Type=oneshot) — it completes before ANY interface is configured,
  so a rotation there has a zero-length window BY CONSTRUCTION, not by being fast.
- R-148: the stub does not need 'fully-up'. stub-first-boot.sh has no pvesh/pct/pveum/qm call
  (grep rc=1); that usage is in felhom-bootstrap.sh under its own After=network-online unit.
  answer.toml.tmpl:27 justifies the current ordering with a dependency that does not exist.
- R-149: the ordering enum has THREE values (before-network, network-online, fully-up), not two.
- MECHANISM (container, not PVE): locked root closes sulogin — 'the root account is locked' for
  both '*' and '!', with a working control. So 'discard' and 'lock' are the SAME outcome for
  recovery, making the escrow decision binary.
- R-150: all four proxmox-first-boot-* units are Condition-guarded; a failed condition is a SKIP,
  so a hook that never ran looks identical to one that succeeded.
- R-151: demo-felhom was installed from an UNCOMMITTED profile — a Tier 0 reference box is not
  reproducible from main.
- Q4: four gates in iso-repack.sh enforce the single-entry menu; default/timeout already settable.

The first mechanism probe was invalid (uid 0 bypassed pam_unix; sulogin had no tty) and a teardown
error (shredding the control plaintext) are both recorded as failures, not massaged.

demo-hp teardown is INCOMPLETE and named as such; the command is recorded, not claimed done.
2026-07-31 12:44:33 +02:00
admin ea00976403 SPIKE: a universal ISO needs a different disk strategy and a locked root
Findings only — no script, profile or build file changed; no ISO built, nothing published.
documentation/audits/SPIKE-universal-iso-2026-07-31.md

- R-139 (HIGH): a disk filter matching >1 device does NOT fail safe. Observed in a nested VM —
  the installer silently picked one of two matching disks and wiped it; validate-answer accepts
  such an answer. The 'filter did not match any devices' guard covers the ZERO-match case only.
- No udev property distinguishes an internal system disk from external media. Measured on
  demo-felhom with its 1TB external attached: ID_BUS='ata' for BOTH, lsblk RM=0 for both, and
  device-info exposes no removability property. demo-hp's NVMe carries no ID_BUS/ID_TYPE at all.
- R-141 (HIGH): the answer schema makes a root credential mandatory, but root-password-hashed='*'
  validates AND installs to completion. [first-boot].ordering accepts 'before-network', the only
  ordering that closes the exposure window structurally.
- Q3: prepare-iso leaves grub.cfg byte-identical to stock (15 entries, automated AND interactive)
  — a two-entry menu is purely a Felhom grub.cfg.tmpl change.
- R-129 resolved: demo-hp's key is the operator's own, added post-install; demo-felhom's IS baked
  by an uncommitted profile.

The reachable-before-rotation measurement FAILED twice and is recorded as failed, not inferred.

Opens R-139..R-147; restates R-128.
2026-07-31 12:12:35 +02:00
admin f6aed82940 host-install v1.22.0 — E-2 Part 2: new boxes get a real backup target, or are told they do not
Every box installed before this got local_backup_target "local" -- the vzdump
target on the SAME physical device as the guest, so a drive failure took the
guest and its only local backup together. E-1 fixed two machines by hand; this
fixes the installer.

Case A: an eligible secondary drive is already mounted -> create felhom-backup on
that drive's own mountpoint via the felhom-backup-target-apply wrapper (create +
grant) and point the primary tier at it.

Case B: system drive only -> the target stays on the system drive and this is
RECORDED AS DEGRADED, not as normal. The install still succeeds: a single-drive
appliance is a valid product, it just cannot survive drive loss.

Phase 0 inverts the emphasis: the installer has NO drive-enrollment step, so on a
fresh appliance Case A almost never fires. The common case is Case B with the
drive arriving later through the wizard (Part 3). Case A covers the reinstall
shape where an agent-generated .mount unit already brings the drive up by fs-UUID.

Eligibility suggests and refuses the absurd, never decides by transport: the
reference backup drive is an external USB HDD and BOTH demo boxes report
removable=0, so a transport rule disqualifies the reference drive and a removable
rule finds no candidate at all.

Scenario F: an already-configured box is never corrected -- an early return plus
setdefault, both load-bearing.

Proofs (installer-logic-tested against extracted functions with stubbed
pvesm/wrapper; NOT install-tested, no reinstall was performed):
  A -> create + grant, resolved felhom-backup
  B -> DEGRADED warnings, resolved local, rc=0 (install not failed)
  F -> skipped, 0 wrapper calls
  F red-proof (guard removed) -> 2 wrapper calls, i.e. it would have "corrected"
      a correct box
2026-07-29 09:06:32 +02:00
admin b4c528801a host-install 1.21.0: F-LEAK — grant FelhomAgentGuest on the scratch VMID band
A failed restore-test's scratch guest never joins the felhom pool, so the pool-scoped
grant cannot reach it and teardown 403s. Ten path-scoped /vms/<id> grants reach exactly
the scratch band and nothing else. Removal path + verify step extended.
2026-07-28 11:05:44 +02:00
Claude Code adf1d1e619 R-82 Slice D/E: installer default 1.20.0 + architecture docs brought current
Slice D.1 — host-install 1.20.0: a FRESH box defaults to local-daily +
offsite-weekly (felhom-pbs, 604800s, keep_last=2). setdefault semantics proven
both ways: fresh gets the tier, an UPGRADE preserves the existing backup block
verbatim — so an in-place upgrade can never silently start writing to an
offsite datastore. Existing boxes are migrated explicitly.

Slice E:
- 07-backup-architecture.md: honest status header per CONTEXT ruling S-2, with
  an explicit STALE-outside-the-PBS-tier verdict (the controller tiers were last
  verified 41 controller versions ago). The PBS row claimed 'PBS on DooPlex'
  (the retired spike store) with no cadence; it now names felhom-pbs ->
  felhom-offsite on ep0 over wg-felhom, weekly, keep_last=2. NOT marked
  ratified — that is Viktor's review of the section 10 list. Discharges R-83.
- 06-offsite-connectivity.md: the target-split remaining-work note collapsed
  (shipped), and records HOW S4.1's tier-aware timeout silently regressed — the
  mechanism was never removed, its INPUT changed when local_backup_target was
  retargeted to 'local'. Also notes S4.1 already diagnosed the teardown 403 as a
  phantom (a timeout consequence, not an ACL gap).
- capability map: new row for recurring offsite backups actually LANDING, as
  distinct from the existing row proving ACTIVATION. IMPLEMENTED, not
  PROVEN-LIVE — the restore round-trip has not completed under the fixed code.
- ROADMAP: R-82 SHIPPED with its remaining gate named, R-83 DISCHARGED, R-84
  left open.
- CONTEXT + REPORT: the arc, including the mid-arc correction I had to make.
2026-07-26 17:54:10 +02:00
admin 485321f694 R-50 Phase A: host-install v1.19.0 island default + hub version sync
- felhom-host-install v1.19.0: portless vmbr9 island bridge, appliance binds
  local_api on 169.254.253.1:8443, writes island_bridge/island_guest_addr,
  pins lan_resolver.host_ip to the LAN IP (Finding-1). --no-island opt-out.
- hub hostInstallVersion 1.16.0 -> 1.19.0 (F-1 sync). hostinstall_gates PASS.
- Pairs with agent v0.96.0 (attaches guest net1). byo unchanged.
Coupling: island install requires agent >= 0.96.0 (vouch first).
2026-07-25 14:17:43 +02:00
admin a2c60edefd docs: golden 0.161.0 published+vouched (R-71 gap CLOSED) + v1.25.0 ISO built; nested drill remains
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NKSN3gSg4TKVBBqkwW2djR
2026-07-23 17:13:28 +02:00
admin b620435afe ISO train v1.25.0 code: OOB belt default appliance leg (F9) + apt no-subscription (Part2) + R-63 Lat2 console font + R-71 golden>=floor build gate
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NKSN3gSg4TKVBBqkwW2djR
2026-07-23 15:38:14 +02:00
admin a12c6f9730 docs(v1.24.0): R-59/R-60/R-61 SHIPPED — CHANGELOG, README, ROADMAP (+R-62), runbook, capability map, drill evidence, REPORT
Virgin-ISO nested drill closed the train: dead-NIC install baked the
fallback (incl. the dead default gateway), the R-59 screen painted
(capture committed beside the spike doc), the cable move healed +
registered at the hub in 23s unaided, and the build's rootpw file
matched the installed box's shadow hash. R-59 SHIPPED with the recorded
deviation (first-boot gate; installer-initrd abort out of scope by
operator ack). R-60 SHIPPED (spike + drill cited; F-P9 route-flush fix
included). R-61 slice 1 SHIPPED. New R-62 row (hub delete-dialog
cosmetics, XS). Capability map: new PROVEN-LIVE row (nested != metal,
said so). Cleanup verified: felhom-pve interfaces byte-identical,
bridge/VMs/ISO removed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
2026-07-22 11:45:55 +02:00
admin af8a7a58a4 docs+scripts: TASK-E parts 4-7 - operator one-pager, rulings, hub build.sh adoption
PROMPT-TEMPLATE: standard 'For the operator' plain-language section, mandatory
for M+ tasks and anything with a STOP.

ROADMAP rulings (operator, 2026-07-21): R-25b full-teardown cascade with three
acks + typed name (M-sized, spec to follow, no longer blocks R-3); R-11 channel
= direct Messenger, doc is the architect's; R-42 option (a), sidecars follow the
app; R-17 delete the archive - spike-lite found NO tooling verb targets it, so
it is an operator console action; R-4 complete (freemail.hu verified).
R-55 + R-41 slice 1 marked shipped; new R-56 (app difficulty classification -
the constructive half of the glance ruling).

scripts/build-hub.sh v1.23.0: the hub build script was outside any repo. Adopted
verbatim + versioned; the build-dir path is now a symlink to it.

felhom-testing skill: the ~1/5 recovery-code 'known flake' is retired - it was a
real defect the test was correctly detecting.
2026-07-21 15:32:56 +02:00
admin ee48a48288 docs: migrate workflow to DooPlex-local execution 2026-07-19 12:16:52 +02:00
admin 13af252907 docs: scripts v1.22.0 + website grid CHANGELOGs, ROADMAP R-38 shipped, R-45/R-46, REPORT
R-38 flips to SHIPPED with the live-canary evidence. Two new items registered
from tonight's work: R-45 (unified async-job feedback — five one-off progress
surfaces now exist and agree on nothing; the two lessons any framework must
encode are "probe the terminal state, never infer it" and "a source reporting
nothing is normal, not broken") and R-46 (verification copies are now visible
but still cannot be browsed and never expire — the invisible-result gap and the
disk-lifecycle gap, designed together because they are the same object).

Pre-invite checklist gains the golden note: build-golden v2.1.0 makes goldens
>= 0.147.x carry all four infra images; no rebuild for this alone.

REPORT-<topic>.md per the parallel-session rule — another session is writing in
this clone tonight, so the shared REPORT.md is left alone.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
2026-07-19 10:00:10 +02:00