Both demo machines are home, unmuted and healthy; one true alarm stands
gates / gates (push) Successful in 24s
gates / gates (push) Successful in 24s
Powered up 2026-08-10 ~09:26 CEST. Both unblocked on the hub, both OK on the approved pair (agent 0.128.0, controller 0.210.0). No false alarm on power-up -- the mute did its job and was removed as the banner said it must be. The hub briefly read "Guests 0/2": the agent's first post-boot report genuinely said stopped, because it caught the guests mid-start. Both corrected to running on the next cycle (07:40:52Z, 07:43:39Z). Transient, not a defect -- confirmed by waiting for the observable rather than assuming it. demo-hp's off-site repository still opens with the box's own credential, 18 snapshots intact including yesterday's rehearsal files, so tonight's 04:15 run has what it needs. demo-felhom's is ORPHANED with no successful run ever; the offsite_stale mail it sent this morning is a TRUE alarm and the remedy is the customer-present recovery ceremony, which needs the operator.
This commit is contained in:
@@ -10,24 +10,22 @@
|
||||
> *Rebuilt from the register on 2026-08-07, from 258 lines. The old "what shipped recently" log is what
|
||||
> the per-repo `CHANGELOG.md` files and the register are for, and is not restated here.*
|
||||
|
||||
## ⚠ BOTH DEMO MACHINES ARE OFF AND MUTED — unmute them when they are home
|
||||
## Both machines are home, unmuted and healthy — one thing still needs you
|
||||
|
||||
**Powered down 2026-08-09 14:08 CEST** for the move back from the vacation home. Guests stopped
|
||||
cleanly first (no vzdump was running, no locks), then the hosts. Confirmed off at the fabric, not
|
||||
merely unreachable: the tailnet is healthy and both peers report *"offline, last seen 1m ago"*.
|
||||
**Back online 2026-08-10 ~09:26 CEST**, both unblocked on the hub, both reporting **OK** on the
|
||||
approved pair (agent 0.128.0, controller 0.210.0). No false alarm fired on power-up. `drill-r50` is
|
||||
untouched and still blocked, as intended.
|
||||
|
||||
**Both customers are BLOCKED on the hub, deliberately, to stop four false alarms an hour into the
|
||||
drive.** Blocking gates every monitor and the notification intake; it does **not** gate config pull or
|
||||
report intake, so the boxes come back normally on power-up.
|
||||
**demo-hp is in good shape.** Its off-site repository still opens with the machine's own key —
|
||||
**18 snapshots, including yesterday's rehearsal files** — so the tier is credentialed and ready; its
|
||||
first scheduled run since the rebuild is tonight at 04:15. Two apps it had before the rehearsal
|
||||
(opengist, privatebin) were never reinstalled; only Calibre-Web was, as the walk needed.
|
||||
|
||||
> **THE TAIL, and it is the reason this banner exists: while they are blocked, a box that FAILS to
|
||||
> come back up is also silent.** When the machines are home and powered on, unblock them and confirm
|
||||
> both report:
|
||||
>
|
||||
> Hub → Customers → **demo-hp** → Unblock, and **demo-felhom** → Unblock.
|
||||
>
|
||||
> Then check both read ONLINE on Hosts. **Until that is done, the hub cannot tell you either box is
|
||||
> in trouble.**
|
||||
**demo-felhom has had no off-site backup for a week, and it will not fix itself.** Its repository is
|
||||
flagged **orphaned**; the last attempt was 2026-08-05 and **no run has ever succeeded**. The hub
|
||||
raised `offsite_stale` this morning and mailed you — that alarm is **true**. The remedy is the
|
||||
recovery ceremony on that machine's own screen, using its recovery code, which is the one thing I
|
||||
should not do without you saying so. *(R-278.)*
|
||||
|
||||
## What works
|
||||
|
||||
|
||||
Reference in New Issue
Block a user