Files
felhom.eu/STATUS.md
T
admin b080ecf411 hub v0.99.0 — the hub can see whether the operator can get in (R-260); G-1 gate closes, R-247 closes
oobDegraded tested five things and the sixth never arrived.

The agent has emitted `operator_key_configured` on every heartbeat since v0.72.0 — the SAME version
that introduced the `oob` stanza carrying it — and store.HostOOBRow mirrored five of the agent's
eight OOB fields. With no field for it, encoding/json discarded the fact on arrival, so a box with
felhom-sshd active, reachable, a valid config and a configured peer reported `ok` with NO OPERATOR
KEY INSTALLED AT ALL. Not a wrong answer: an answer to a question nobody was asking.
`operator_peer_configured`, which the hub did read, only says the peer IP is in desired-state — that
OOB is MEANT to work, not that entry is possible.

Now decoded: operator_key_configured, plus wg_handshake_age_s and healed_at. The last two ride the
ALERT TEXT and are deliberately NOT in the predicate — widening a check beyond the fact that is now
arriving is how a check stops being read.

SCENARIO F, decided on a measurement rather than a preference. operator_key_configured decodes as a
POINTER: nil = the agent never said, reported distinctly and never as ok. The version gate was
rejected because the field and its stanza shipped in the SAME agent version (v0.72.0), so a stanza
without the field cannot come from any released agent; the fleet is 0.113.0/0.127.0 and the vouched
floor is 0.127.0. Handled explicitly anyway and pinned, because "cannot happen" is a claim this
project has been burned by.

THE MESSAGE NAMES THE FAULT. oobDegradedReason is the single source for both predicate and text, so
the alert can never name a different fault from the one that fired. The old form derived it
separately and had a vocabulary of two — unreachable, or config invalid — with no way to say the key
is missing. The operator reads this at 07:00.

TESTS DRIVE THE DECODE BOUNDARY. Every hub OOB test before this built a HostOOBRow by hand, and a
test written that way CANNOT SEE A FIELD THAT NEVER DECODES — which is how this held a green suite
for five weeks. The pre-existing fixture oobReport() also omitted the field, so those scenarios ran
against a report shape no released agent produces (same family as R-262). Both fixed.

Red-proofs, 8 expected outcomes and 0 wrong, each with the mutation asserted applied: dropping the
field returns the false ok; an unconditional check alerts a healthy box; unknown-as-ok restores the
silent pass.

G-1 CLOSED — scripts/wire_contract_gate.py shipped as ranked, built BEFORE the fixes and seen
failing on 40 fields (documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md). Two instrument
defects the control caught first: a substring false negative (grep -F healed_at matched
privsep_healed_at) and treating dr_recipe as wholly opaque when its top-level sections ARE decoded
through an allow-list that already cost offsite_restic (R-122).

The prompt for this session said "465 emitted tags, eight unreachable". Checked against the repo:
R-260 said "at least eight DECISION-BEARING facts", never eight tags. The real count is 40.

R-260 CLOSED (class gated, sharpest instance fixed). R-247 CLOSED (controller v0.209.0). R-264
MINTED and OPEN — the 21 facts with no consumer, allowlisted with reasons so that gating the class
could not be mistaken for deciding them. Still open and named: R-246, R-255..R-259, R-261..R-263,
and C7's test-comment half.

Capability map checked: it claims OOB access is implemented, never monitored, so no row was untrue;
what was untrue sat one layer down and the row now records it.

repo_gates --fast: all 8 OK. go build/vet/test green in hub, run separately from this commit.
2026-08-08 08:47:02 +02:00

94 lines
6.0 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-08-08.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. Not `CONTEXT.md`, which is technical
> state written for Claude Code. **Items, not paragraphs. One screen.** If it does not fit, something
> belongs in the register instead.
>
> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what
> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.*
## What works
A blank machine boots the Felhom disc, installs itself unattended, and is claimed by the customer, who
sets their own password. They install apps from a catalogue of fifty-three, share files over the home
network, and open apps from a launcher or a shared link. Backups run on their own to three places — the
machine's drive, a second drive, and an encrypted off-site copy.
**The backup promise is proved, and so is getting the data back yourself.** A machine has been
destroyed on purpose and its files came back byte for byte identical — four times now, including a
filename with Hungarian accents. On **2026-08-07 the household's own journey passed for the first
time**: someone with a browser and their recovery code got everything back with **no command line
inside the machine at any point**. From logging in to seeing what is in the store took **72 seconds**.
*(R-201 — closed.)*
**And the two rough edges that walk found are gone** — after a rebuild the restore used to stop dead
twice; both refusals now say what happened, that nothing is lost, and link to the screen that fixes
it. *(R-252, R-253 — closed 2026-08-08.)*
## What's broken
- **Nothing new is broken.** All three secret-in-page faults are fixed; what remains is that the
*check* against a fourth covers 4 pages of 27, and the cheap one covering all of them is blind to
the shape that actually shipped. *(R-255)*
- **The machine's own screen keeps telling an already-paired box to pair itself** — 25 minutes after it
was paired, on a screen that promises it refreshes itself. *(R-214, R-235)*
- **A rebuilt machine cannot create a new recovery code at all.** *(R-221)*
- **A backup that covered nothing still calls itself „Sikeres".** The state is honest; the word is not.
*(R-240)*
- **A machine waiting for its recovery code can stop backing up off-site without alarming us.** After
a *rebuild* we ARE told; the gap is a box reaching that state with no working tier behind it. *(R-243)*
- **The card offering to reopen set-aside backups promises more than we can deliver** — we keep the old
sealed package, but nothing can open it. *(R-202)*
- **Deleting a customer leaves rows behind** on every test machine ever torn down, while reporting a
clean teardown. No secrets involved, but it accumulates with each walk. *(R-244)*
- **Putting restored files back where they belong is still a manual step.** *(R-213)*
## Fixed today — the thrown-away sentence, and a check so there is no next one
We could not tell whether your engineer could get into a machine. The machine says so every few
minutes; **the hub had nowhere to put the sentence and discarded it on arrival**, so a box with the
door open, the lock working and **no key issued** was reported as fine. Not a wrong answer — an answer
to a question nobody was asking. Fixed, and the alert now **names the missing key** instead of saying
"access degraded". *(R-260, R-247 — closed; hub v0.99.0, controller v0.209.0.)*
**The check was built first and watched failing on 40 facts, before a single one was fixed** — the
night before, an off-the-shelf tool for a neighbouring shape was rejected for failing exactly that
test. Of the 40: three now change what we are told, sixteen are genuinely redundant, and **twenty-one
are recorded as undecided rather than quietly waved through** *(R-264)* — the strongest being
per-guest network health, which we already lost 1 h 15 m to once.
## What we're working on
- **Widening the check** so a fourth secret-in-a-page is caught by a machine, not by someone. *(R-255)*
- **Deciding the twenty-one** — for each: give it a reader, or stop sending it. *(R-264)*
- **Proving the hub really keeps the old sealed key** when a machine re-seals. *(R-198)*
- Still open from the overnight sweep, none urgent: *(R-256…R-259, R-261…R-263)*
## Waiting on you
- **One approval: the new base image.** Tonight's release is baked, published and byte-checked, and
**installations still receive yesterday's version until you press Save.** Hub → Configuration →
Day-0 artifacts → Golden **0.208.0** → Save. One field moves; the other two are already right and
were checked. Reversible — re-select 0.207.0 and Save. *(R-242)*
- **Third time in three days, so worth a minute.** A check now catches the *baking* being forgotten;
**nothing catches the approval being forgotten** — tonight's bake proved it, going green before the
approval existed. Two ways to close it are written up, neither built. *(ROADMAP G-8)*
## DooPlex infrastructure — separate from the product
*Kept under its own heading rather than dropped: these are real asks that need you, but they concern
the machine all this is built on, not what a customer receives. Mixing them in is why the page stopped
being readable.*
- **DooPlex's own backup keeps every copy inside the same box, and is silent when it fails.** *(R-232)*
- **193 old images exist only on this machine**, ~27 GB against 199 GB free — clutter, not space. *(R-210)*
- **The hub password needs rotating** — a diagnostic printed it into a session log; nothing suggests
anyone else saw it. *(R-132)*
- **One thing to read after DooPlex next restarts** — the second-SSD move has never survived a reboot;
it writes PASS/FAIL to `/var/log/felhom-store-postboot-check.log`. On PASS, 34 GB comes back. *(R-209a)*
- **Backup scripts on DooPlex are unversioned host state** *(R-231)*, and the instruction-file
follow-ups each need a decision rather than an edit *(R-229, R-230)*.