hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect).
This commit is contained in:
@@ -476,7 +476,14 @@ Report MUST include:
|
||||
6. **Deployed versions** + `docker ps` / agent-service / hub-pod verification output.
|
||||
7. **NOT yet live-validated — awaiting supervised [Bx]:** explicit list (the real
|
||||
put-data → operate → integrity flow).
|
||||
8. **Teardown evidence — all three §13 layers**, if the run provisioned anything: the machine deleted,
|
||||
8. **Evidence copied off BEFORE each revert** — for every phase that ran on a machine, the logs were
|
||||
pulled to the evidence directory **at the end of that phase**, before any revert, snapshot restore
|
||||
or teardown, **including the intermediate ones**. *The intermediate revert is the one that gets
|
||||
forgotten: two sessions lost a phase's logs to a mid-run revert to `virgin` on 2026-08-12 and
|
||||
2026-08-13 — same machine, same point, three days apart (R-320).* **If a phase's evidence is
|
||||
already gone, the report says so plainly and the finding is REPRODUCED independently** — that is the
|
||||
expectation, not an improvisation. A quotation read live and no longer re-readable is named as such.
|
||||
9. **Teardown evidence — all three §13 layers**, if the run provisioned anything: the machine deleted,
|
||||
`pvesm status` before/after with the space returned, and the **hub-side record's disposition named**
|
||||
(deleted / retained-with-reason / gate-blocked-with-the-command). A run that provisioned nothing says
|
||||
so. "Teardown clean" without layer 3 is not a report — it is the `sess-c` failure.
|
||||
|
||||
@@ -126,6 +126,7 @@
|
||||
| Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 |
|
||||
| WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` |
|
||||
| The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe |
|
||||
| The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing |
|
||||
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → **R-133** |
|
||||
|
||||
## F. Notifications & monitoring
|
||||
|
||||
@@ -0,0 +1,170 @@
|
||||
# REPORT — The removal that works once, and a status page that says what it is asking for (2026-08-13)
|
||||
|
||||
**Shipped:** `installer-v1.28.0`, published and verified against the live URL.
|
||||
**Register:** R-305 CLOSED (by R-316), **R-316 / R-317 / R-318 opened**; ceiling R-315 → **R-318**.
|
||||
**Venue:** `drill-r50` only. Neither demo box reinstalled. `peti-felhom` not contacted. Nothing deleted.
|
||||
|
||||
---
|
||||
|
||||
## 1. Cycle 3, before and after
|
||||
|
||||
**Before — on the PUBLISHED v1.27.0, from `virgin`, exit 1:**
|
||||
|
||||
```
|
||||
[ERROR] a resolver is already bound to :53 on this host:
|
||||
udp UNCONN 0 0 0.0.0.0:53 … users:(("dnsmasq",pid=7076,fd=4)) …
|
||||
[ERROR] a resolver is already bound to :53 on this host — Felhom needs the guest reachable by name on your LAN.
|
||||
Stop or reconfigure that resolver, OR point your LAN DNS at the guest's address, then re-run.
|
||||
(Felhom does NOT touch DNS services on a host it does not own — this is a refusal, not a change.)
|
||||
THIS LOOKS LIKE OURS. A previous Felhom install leaves the dnsmasq PACKAGE installed and its unit
|
||||
enabled (only our config snippet is removed), and unconstrained it binds 0.0.0.0:53 — which is what
|
||||
this gate is seeing. If this host had no dnsmasq before Felhom, clear it with:
|
||||
systemctl disable --now dnsmasq
|
||||
Then re-run this installer. If dnsmasq is YOURS, leave it and use one of the two routes above.
|
||||
[ERROR] PRE-FLIGHT FAIL (exit 1) — fix the finding above and re-run
|
||||
```
|
||||
|
||||
**After — v1.28.0, three fresh cycles from `virgin`, exit 0:**
|
||||
|
||||
```
|
||||
:53 now: 0
|
||||
CYCLE 3 rc=0
|
||||
[INFO] host DNS (:53): free
|
||||
[OK] pre-flight passed
|
||||
[OK] PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed
|
||||
```
|
||||
|
||||
The three cycles were run on bytes **sha256-identical** to what the live URL now serves
|
||||
(`cca1dedd9b7c152c…`, compared three ways: live URL, tested file, repo main).
|
||||
|
||||
## 2. The mechanism, at `file:line`
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| ownership **recorded** | `felhom-host-install.sh:1718` — preflight, `dpkg-query -W -f='${Status}' dnsmasq \| grep -q "install ok installed"` |
|
||||
| ownership **read** | `:1097` (`_state_get dnsmasq_preexisting`) at uninstall |
|
||||
| state file **deleted** | `:1268` — *after* the read. **The order was already correct.** |
|
||||
| Felhom's dnsmasq snippets removed | `:1162`, in the loop just above the ownership decision |
|
||||
|
||||
**Why cycle 2 concludes "pre-existing": package presence alone.** The preflight asks dpkg one question
|
||||
and nothing else — it does **not** consult the absence of a record, and it cannot, because the record
|
||||
was deleted with the state file. v1.27.0 stopped the unit and left the package, so the answer stayed
|
||||
`yes` and our own package became "the household's" one cycle later.
|
||||
|
||||
**What the first uninstall did and did not remove:** snippets — removed. Unit — stopped + disabled.
|
||||
State file (and with it the record) — removed. **Package — left.** That last one is the whole defect.
|
||||
|
||||
## 3. What v1.28.0 changes
|
||||
|
||||
When the record says we installed it, the uninstall **removes the package** as well as stopping the
|
||||
unit. Order unchanged: read the record → act → delete the state.
|
||||
|
||||
**Two packages are recorded, not one.** `dnsmasq` ships the systemd unit; **`dnsmasq-base` ships
|
||||
`/usr/sbin/dnsmasq`** (`dpkg -S`, measured on the box). Separately installable, so each is recorded at
|
||||
preflight and taken back only if we added it.
|
||||
|
||||
**Guard rails:** ownership **read, never inferred**; the dependency check is an **`apt-get -s purge`
|
||||
simulation** that proceeds only if the removal set is a subset of ours, else stop+disable **naming the
|
||||
blocking package**; never interactive; **never fatal** — a wedged apt is recorded and restated in the
|
||||
closing NOTE; and the success is **re-queried** rather than read off apt's exit code.
|
||||
|
||||
## 4. Scenarios and red-proofs
|
||||
|
||||
| Scenario | Result |
|
||||
|---|---|
|
||||
| **A** three cycles, fixed | install 3 **PASSES**, `:53` free at every uninstall |
|
||||
| **B** resolver pre-dates Felhom | **untouched** — `install ok installed`, unit active, both packages recorded `yes` |
|
||||
| **C** something depends on it | **not purged**; log named `household-dns-thing`; stop+disable; `:53` free; dependent survived |
|
||||
| **D** no ownership record | **untouched**, reason logged, exact command named |
|
||||
|
||||
**Red-proofs — mutation asserted applied before each run:**
|
||||
|
||||
| Mutation | Outcome |
|
||||
|---|---|
|
||||
| remove the purge call | cycle 2 records `yes`; **cycle 3 refuses, exit 1, in those exact words** |
|
||||
| remove the ownership check | **the household's resolver is PURGED** (`unknown ok not-installed`) |
|
||||
| infer ownership when no record | **the guess is taken** — a field box loses its own DNS |
|
||||
|
||||
**One honest note on red-proof B.** Its first run appeared to pass *for the wrong reason*: apt failed
|
||||
with `dpkg was interrupted` (a broken state my own earlier manual package juggling left), so the
|
||||
resolver survived by accident rather than by the guard. I repaired dpkg and re-ran, and it then failed
|
||||
as required. Worth stating twice over: the mutation looked like a passing guard, and **my own fix's
|
||||
apt-failure path was incidentally observed doing exactly what it should** — logging, continuing, naming
|
||||
the command.
|
||||
|
||||
## 5. The machines already in the field
|
||||
|
||||
**Measured, not reasoned.** A field box is one with the package installed and **no record** — scenario
|
||||
D. Its uninstall leaves the resolver running with the reason logged, and its next install refuses with
|
||||
the message quoted in §1.
|
||||
|
||||
**Is there an honest durable marker? No, and none can be invented.** The Felhom `/etc/dnsmasq.d/
|
||||
felhom-*.conf` snippets are deleted by the uninstall's own loop *before* the ownership decision; the
|
||||
state file carrying the record is deleted at `:1268`; nothing under `/etc/felhom*` survives.
|
||||
`/var/log/dpkg.log` does record the install and is a **timestamp** — refused by the standing rule as a
|
||||
heuristic dressed as a fact. → **R-318**
|
||||
|
||||
**The preflight message, judged as a customer would:** it states the finding, keeps its two routes and
|
||||
its written promise not to touch DNS on a host we do not own, and adds the one thing that gets a person
|
||||
moving — *"THIS LOOKS LIKE OURS"* plus one exact command. It hedges correctly (*looks like*). Its
|
||||
weakness is that it asks them *"did this host have dnsmasq before Felhom?"* — precisely the question we
|
||||
can no longer answer for them. **That is a mechanism, not a rule: the command is on the screen at the
|
||||
moment it is needed.**
|
||||
|
||||
## 6. The defect I nearly shipped
|
||||
|
||||
The agent decides whether to install dnsmasq with `os.Stat("/usr/sbin/dnsmasq")`
|
||||
(`felhom-agent/internal/lanresolver/lanresolver.go:105`) — but that path belongs to **`dnsmasq-base`**,
|
||||
while the unit comes from **`dnsmasq`**. Purging only `dnsmasq` would leave the binary, so the next
|
||||
install would skip the apt step and then fail to enable a unit that is gone — **a silent resolver where
|
||||
today there is at least a visible refusal.** That is why ownership is recorded per package.
|
||||
|
||||
It remains reachable where `dnsmasq-base` pre-dated Felhom (we correctly keep it). **Pre-existing, not
|
||||
introduced here**, filed as **R-317**, and the uninstall now says so out loud rather than leaving it to
|
||||
be found from a resolver that never came up. Fixing it is a one-line agent change, deliberately **not**
|
||||
made here to keep this session to one repo.
|
||||
|
||||
## 7. Publication
|
||||
|
||||
`installer-v1.28.0` tagged; **both** `--ref`s in `manifests/webpage.yaml` moved (sidecar 327, init 372);
|
||||
ArgoCD synced to `823cd29`; live deployment carries the new ref.
|
||||
|
||||
```
|
||||
live URL serves: SCRIPT_VERSION="1.28.0" (was 1.27.0)
|
||||
_dnsmasq_purge_owned present in the served file: 3
|
||||
```
|
||||
|
||||
**Pushing publishes nothing here** — `/scripts/` follows the tag. The runbook still says otherwise;
|
||||
**R-309 remains open and is not forgotten**, deliberately out of scope.
|
||||
|
||||
## 8. Teardown — four layers
|
||||
|
||||
| Layer | State |
|
||||
|---|---|
|
||||
| **machine** | `drill-r50`: 0 guests, no `/var/lib/felhom-install`, dnsmasq `not-installed` at the end |
|
||||
| **host** | `drill-r50` **reverted to snapshot `virgin`, powered off**; `qemu.pid` removed. DooPlex was drill host only |
|
||||
| **hub** | `drill-r50-0a4f9a` **RETAINED** — pre-existing since 2026-07-25, re-used not duplicated. **No new host or customer row was created this session.** Nothing deleted |
|
||||
| **off-site** | **Not contacted at all** — no restic, sftp or endpoint call was made. `demo-felhom`: `last_status=ok`, `escrow=escrowed`, **no abandon countdown** |
|
||||
|
||||
Both demo boxes untouched and healthy: controller **0.214.0**, agent **0.129.0**, `Up … (healthy)`.
|
||||
|
||||
## 9. Evidence, and a repeated miss
|
||||
|
||||
Logs at `documentation/audits/evidence-r316-2026-08-13/`: the fixed cycles (`F1`, `F2`, `F3`), all four
|
||||
scenarios (`B`, `C`, `D`), and every red-proof (`RPA*`, `RPB*`, `RPD*`).
|
||||
|
||||
**The Part 1 logs did not survive.** They were on the drill VM's disk and were destroyed by the revert
|
||||
to `virgin` between Part 1 and Part 2 — **the same mistake as Tuesday, in the same place.** The §1
|
||||
quotation is verbatim from the live run as read at the time, and `RPA1/RPA2/RPA3.log` are an
|
||||
independent reproduction of the identical three-cycle failure, retained. Recorded rather than glossed;
|
||||
the fix is procedural and I have now got it wrong twice.
|
||||
|
||||
## 10. Observations — noticed, not acted on
|
||||
|
||||
- **`dnsmasq-base` is flagged "automatically installed and no longer required"** after `dnsmasq` goes.
|
||||
We deliberately do not `autoremove` — that would be a blast radius nobody asked for.
|
||||
- **`apt-get -s purge` exits 0 even when it prints `E: dpkg was interrupted`.** The simulation output is
|
||||
the signal, never the exit code — which is why the guard parses `Remv`/`Purg` lines rather than
|
||||
trusting `$?`. Another entry for the exit-codes-that-lie class.
|
||||
- The drill VM's `pveam` index is stale on `virgin` and needs `pveam update` before listing templates —
|
||||
already recorded yesterday, hit again today in passing.
|
||||
@@ -183,6 +183,14 @@ no first-boot hook at all, so this question now governs **operator-built images
|
||||
`main` on a 30-second period (`documentation/.../day0-install.md:150-152`) — no release tag, no
|
||||
staging copy, no version selector. Pushing the script publishes it.
|
||||
|
||||
> **⚠ SUPERSEDED 2026-08-03, annotated 2026-08-13 (R-110, R-309). True on the day it was written;
|
||||
> false ever since.** `/scripts/` is now served from the tag `installer-v<SCRIPT_VERSION>` and
|
||||
> **pushing publishes nothing** — publication is a tag plus **two** `--ref` pins in
|
||||
> `manifests/webpage.yaml`. The dated finding is kept as written rather than rewritten, because it
|
||||
> is a record of what was true then; **the note is here because this paragraph cited
|
||||
> `day0-install.md` as its source, which is how the wrong sentence spread.** Current procedure and
|
||||
> the outside-verification command: `runbooks/day0-install.md` §C.1.
|
||||
|
||||
3. **Console display does not exist.** `felhom-host-install.sh` does not write `/etc/issue`,
|
||||
`/etc/issue.net` or any MOTD (grep: no match). PVE's own `/etc/issue` banner is what the console
|
||||
shows after install — observed in this session's screendumps as
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -26,6 +26,13 @@
|
||||
> **Point of no return:** booting the install-armed stick (the auto-installer needs no confirm).
|
||||
> Recovery unchanged: stock ISO + `felhom-host-install.sh` — and since RESET is the plan, there
|
||||
> is nothing to restore. Demo-felhom.eu is down for the window (~2 h).
|
||||
>
|
||||
> **EVIDENCE RULE — standing rule 5 (R-320), and it applies to every phase boundary below.** The logs
|
||||
> of a phase are copied OFF the machine at the end of **that** phase, before any revert, snapshot
|
||||
> restore or teardown — **including the intermediate ones, which are the ones that get forgotten.**
|
||||
> Two sessions lost a phase's logs exactly this way on 2026-08-12 and 2026-08-13, same machine, same
|
||||
> point. Make the pull the last act of the phase. **If evidence is already gone: say so plainly and
|
||||
> reproduce the finding independently** — do not let a lost log quietly become a softer claim.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -167,9 +167,44 @@ chmod +x felhom-host-install.sh
|
||||
./felhom-host-install.sh -h | head -3 # sanity: must print v1.15.0 (or newer) — DR-tier-by-default ships the full plumbing
|
||||
```
|
||||
|
||||
This URL is the website's git-sync working tree tracking `main` on a 30-second period
|
||||
(`manifests/webpage.yaml`) — it is always the current `main` script. There is no release tag, no
|
||||
staging copy and no version selector; pushing `scripts/felhom-host-install.sh` publishes it.
|
||||
**⚠ THIS PARAGRAPH USED TO SAY THE OPPOSITE OF THE TRUTH, and two sessions were misled by it before
|
||||
it was corrected on 2026-08-13 (R-309, R-110).** The old text read *"it is always the current `main`
|
||||
script … pushing `scripts/felhom-host-install.sh` publishes it."* **Pushing publishes NOTHING, and has
|
||||
not since R-110 shipped on 2026-08-03.** Believing the old sentence is dangerous in **both**
|
||||
directions: it makes an operator think a pushed fix is live when it is not, and think a pushed mistake
|
||||
is live when it is not. (Measured 2026-08-12: this URL served `1.25.0` while `main` held `1.27.0`,
|
||||
three and a half hours after the push.)
|
||||
|
||||
`manifests/webpage.yaml` runs **two** git-syncs against **different refs**: the website tracks `main`,
|
||||
and `/scripts/` is checked out from the tag **`installer-v<SCRIPT_VERSION>`**. So this URL serves the
|
||||
**tagged** installer, not `main`.
|
||||
|
||||
**What actually publishes it — three acts, and the middle one is two lines, not one:**
|
||||
|
||||
1. Bump `SCRIPT_VERSION` in `scripts/felhom-host-install.sh` (the single version source) and push to
|
||||
`main`. *Nothing is published yet.*
|
||||
2. Cut and push the tag `installer-v<new SCRIPT_VERSION>`.
|
||||
3. Move **BOTH** `--ref=installer-v…` pins in `manifests/webpage.yaml` — the git-sync **sidecar** and
|
||||
the **init container** (today lines 327 and 372) — commit, and sync ArgoCD. **Moving one pin is the
|
||||
trap**: the running pod keeps serving until it restarts, and then a fresh pod seeded by the stale
|
||||
init container serves the OLD script with no error anywhere.
|
||||
|
||||
**To roll back:** move the tag back and wait ~30 s. No ArgoCD sync, no deploy — that is the emergency
|
||||
lever; fix forward afterwards. **Do not pin the website to the tag**, or every copy edit becomes a
|
||||
release. `hostinstall_gates.py` gate 6 fails if the manifest stops naming an `installer-v…` tag or if
|
||||
the website stops tracking `main`.
|
||||
|
||||
**How to verify from OUTSIDE that the published version actually changed** — a push, a green sync and a
|
||||
correct-looking manifest are each consistent with nothing having been published, so ask the public URL
|
||||
rather than the repository:
|
||||
|
||||
```bash
|
||||
curl -fsS https://felhom.eu/scripts/felhom-host-install.sh | grep -m1 '^SCRIPT_VERSION'
|
||||
```
|
||||
|
||||
It must print the **new** version. Confirmed by this method on 2026-08-13: served `1.28.0`, `main`
|
||||
`1.28.0`, both manifest pins `installer-v1.28.0` — the three agreeing is the observation, and any one
|
||||
of them alone is not.
|
||||
|
||||
### C.2 Preview (recommended)
|
||||
|
||||
|
||||
@@ -25,6 +25,25 @@ Fences name **acts**, not machines. "Do not destroy demo-hp's `drill-r50` fixtur
|
||||
demo-hp to host a throwaway VM" are unrelated; only the first has ever been meant. Read a per-machine
|
||||
prohibition as covering the act it names and nothing more.
|
||||
|
||||
## Before you revert it — take the evidence off first
|
||||
|
||||
> **A phase's evidence is copied off the machine at the end of THAT phase, before any revert, snapshot
|
||||
> restore or teardown. Not at the end of the session.** (Standing rule 5, R-320.)
|
||||
>
|
||||
> **The intermediate revert is the one that gets forgotten.** Both losses this project has recorded
|
||||
> were the *middle* teardown, never the final one — the Phase A logs of the 2026-08-12 retained-key
|
||||
> drill and the Part 1 logs of the 2026-08-13 R-316 run, both on `drill-r50`, both destroyed by a
|
||||
> revert to `virgin` between phases, three days apart. Both times the conclusions survived only
|
||||
> because the quotations had been read live and an independent reproduction happened to exist. **That
|
||||
> is luck.**
|
||||
>
|
||||
> **The mechanism:** make the pull the last act of the phase, not a step to remember later —
|
||||
> `pct pull` / `scp` into `documentation/audits/evidence-<topic>-<date>/` on DooPlex, then revert.
|
||||
> Tier 0 machines are *disposable*, which is precisely why nothing you need may be left on one.
|
||||
>
|
||||
> **If you notice it is already gone:** say so plainly in the report and **reproduce the finding
|
||||
> independently**. That is the documented expectation, not an improvisation.
|
||||
|
||||
| Tier | Meaning | Machines |
|
||||
|---|---|---|
|
||||
| **0 — disposable. Reach here first.** | Exists to be broken; reinstalling is a routine afternoon, not an incident. **A drill that needs a victim uses one of these.** | `demo-hp` (t740), `demo-felhom` (N100) |
|
||||
|
||||
@@ -56,6 +56,15 @@ roles. **A file being open in the editor is NOT an instruction. If no task is st
|
||||
"working" and "stopped entirely".
|
||||
4. **A recommendation that is not followed gets one line saying why.** Silence reads as agreement and
|
||||
the disagreement is lost.
|
||||
5. **Evidence is copied off the machine at the end of the phase that produced it — before any revert,
|
||||
snapshot restore or teardown. Not at the end of the session.** The intermediate revert is the one
|
||||
that gets forgotten; both losses were the *middle* teardown, never the final one. **The mechanism,
|
||||
because a rule without one is a wish:** the last act of a phase that ran on a machine is `scp`/`pct
|
||||
pull` its logs to the evidence directory on DooPlex — the same act that ends the phase, not a
|
||||
separate step to remember later. **And when a session notices the evidence is already gone: say so
|
||||
plainly in the report and REPRODUCE it independently.** That is the documented expectation, not an
|
||||
improvisation to be invented under pressure — it is what both sessions did, and it is the only
|
||||
reason two sets of conclusions survived.
|
||||
|
||||
<!--
|
||||
R-96 incident record (committed 2026-07-27) — rationale, not directives.
|
||||
@@ -72,6 +81,16 @@ R-96 incident record (committed 2026-07-27) — rationale, not directives.
|
||||
4. Twice in the R-88/R-97 arc a review point was absorbed rather than argued: R-84 was folded into
|
||||
R-82 without a word, and R-97a's operator-only guard was dropped while the claim it was meant to
|
||||
enforce got committed as a comment — which is how a false guarantee shipped and survived a release.
|
||||
5. Twice in three days, on the SAME machine and at the SAME point. 2026-08-12, the retained-key drill:
|
||||
the Phase A logs lived on `drill-r50`'s disk and were destroyed by the revert to `virgin` between
|
||||
Phase A and Phase B (`audits/DRILL-retained-key-2026-08-12.md` §11.5). 2026-08-13, R-316: the Part 1
|
||||
logs, same disk, same revert, between Part 1 and Part 2 (`audits/REPORT-r316-installer-v1.28.0-2026-08-13.md`
|
||||
§9 — "the same mistake as Tuesday, in the same place"; preserved out of `REPORT.md`, which is
|
||||
overwritten every session). Both times the golden-bake runbook's existing "scp the log OUT first"
|
||||
was applied to the FINAL teardown and not the intermediate one. Both times the conclusions survived
|
||||
only because the quotations had been read live and an independent reproduction happened to exist —
|
||||
that is luck, and the second occurrence is what makes it procedural rather than another apology.
|
||||
Filed as R-320.
|
||||
-->
|
||||
|
||||
## Shared conventions
|
||||
|
||||
Reference in New Issue
Block a user