R-379: --clear-restore-hold now states the required restart
gates / gates (push) Successful in 11s

It runs as a second process: it clears settings.json but the running controller
keeps its in-memory copy and goes on refusing. Measured on demo-hp - clear
succeeded, file correct, start button still refused until a restart.

Also records the lost-update window between the two processes, and why clearing
through the running controller (the right shape) needs an operator tier the
controller's HTTP surface does not have.
This commit is contained in:
2026-08-22 18:21:47 +02:00
parent 5b52a5964d
commit 1a2405e86f
2 changed files with 33 additions and 1 deletions
+17
View File
@@ -1,3 +1,20 @@
## v0.220.2 — the operator route to clear a hold now says the restart is required (2026-08-22, R-379)
**MinAgent: 0.129.0** (unchanged)
`--clear-restore-hold` runs as a SECOND process. It clears the hold in `settings.json`, but the
RUNNING controller holds its own in-memory `Settings` and keeps refusing until it reloads. Measured on
`demo-hp`: the clear succeeded, the file was correct, and the customer's start button still refused —
until the controller was restarted, after which the app started normally.
The command now prints the restart it needs. **A route the operator believes worked, and did not, is
worse than no route**, and this was found by using it rather than by reading it.
Recorded with it: there is a lost-update window while both processes hold the file. Restarting
promptly closes it. Clearing through the running controller would remove both problems and is the
right shape later; it needs an operator tier the controller's HTTP surface does not have today — it
authenticates as the customer, and a customer clearing their own hold is what the hold exists to
prevent.
## v0.220.1 — the rollback poured the undo into a container that no longer existed (2026-08-22, R-379) ## v0.220.1 — the rollback poured the undo into a container that no longer existed (2026-08-22, R-379)
**MinAgent: 0.129.0** (unchanged) **MinAgent: 0.129.0** (unchanged)
+16 -1
View File
@@ -200,7 +200,22 @@ func main() {
fmt.Fprintf(os.Stderr, "no restore hold in force for %q — nothing was changed\n", *clearRestoreHold) fmt.Fprintf(os.Stderr, "no restore hold in force for %q — nothing was changed\n", *clearRestoreHold)
os.Exit(2) os.Exit(2)
} }
fmt.Printf("restore hold cleared for %s — the app may be started again. Check its data first: the undo copies are in its unit's db-dumps dir.\n", *clearRestoreHold) // THE RESTART IS NOT OPTIONAL, and saying so is the difference between a route that works
// and one an operator believes worked. This runs as a SECOND process: it clears the hold in
// settings.json, but the RUNNING controller holds its own in-memory Settings and keeps
// refusing until it reloads. Measured on demo-hp 2026-08-22 — the clear succeeded, the file
// was correct, and the customer's start button still refused until the controller restarted.
//
// There is also a lost-update window while both processes hold the file: if the running
// controller saves settings.json from its own copy before the restart, the clear is undone.
// Restarting promptly closes it. A route that clears through the running controller would
// remove both problems and is the right shape later; it needs an operator tier the
// controller's HTTP surface does not have today (it authenticates as the CUSTOMER, and a
// customer clearing their own hold is precisely what the hold exists to prevent).
fmt.Printf("restore hold cleared for %s in settings.json.\n", *clearRestoreHold)
fmt.Printf(" NOW RESTART THE CONTROLLER, or it will keep refusing to start the app:\n")
fmt.Printf(" systemctl restart felhom-controller-bootstrap.service\n")
fmt.Printf(" Check the app's data first — the undo copies are in its unit's db-dumps dir.\n")
os.Exit(0) os.Exit(0)
} }