Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure. Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/). Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the runbook, STATUS, CONTEXT, R-468 and the gate docstring. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,5 +1,10 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-13 (third pass) — the Update button now takes a backup first and tells the truth. It
|
||||
is live on both machines and I walked every case on the HP, including putting a broken update back
|
||||
from its backup. TWO THINGS NEED YOU: item 15 (the demo machines no longer get new releases by
|
||||
themselves between golden images) and item 16 (apps with no second-drive copy cannot be updated).**
|
||||
|
||||
**Updated 2026-09-13 (second pass) — you decided both open items. The database engine now finishes
|
||||
its own conversion on the four MariaDB apps; the upgrade machine proved it and it landed on the HP
|
||||
without a ripple. Goldens are now weekly and before any install, not per release, and the gate
|
||||
@@ -230,6 +235,49 @@ nothing.*
|
||||
|
||||
13. **"Delete my data too" now deletes the data — or tells you it could not.** Until today, when a customer removed an app and ticked the box, the box said it worked and left everything on the drive (128 MB of a Nextcloud on 2026-09-01). The cause: the removal asked one global setting for the drive, and no machine fills that setting in. Every other part of the box already asks the app itself where its data is. Now the removal does too. If the box cannot work out where the data is, it refuses and keeps the app, so you can try again — it never again reports success over data left behind. Proven on the HP with a throwaway Nextcloud: 63 MB the app wrote itself was gone after removal, and the answer listed it; the refusal was shown with the app still in place; an app with no drive data gets a plain "nothing to delete" note. Live on both machines (0.236.0). No standing app was touched. **If you do nothing:** nothing to do.
|
||||
|
||||
14. **The Update button now takes a backup first, and tells the truth.** Nothing needs you.
|
||||
Before today, pressing Update pulled the new version at once, checked nothing, and said
|
||||
"done" while the app could already be crashing. Now:
|
||||
- It refuses first if it must: the app is held, a backup is running, or memory or disk is short.
|
||||
It also refuses if the app has no backup it could be put back from.
|
||||
- If the backup is older than a day, it makes a fresh one first.
|
||||
- It waits for the app to actually be healthy before it says it worked. The button shows each step.
|
||||
- If the new version does not come up, the app is stopped and held. The page names the backup to
|
||||
restore it from. The box never puts the old version back by itself, because we measured that
|
||||
this works for some apps and breaks others.
|
||||
|
||||
**Proven on the HP with a throwaway app:** a real upgrade, a stale backup, a version that does not
|
||||
exist, a version that never starts, and the restore from the named backup back to the old version.
|
||||
**One problem showed up during the test and is fixed:** while an update was waiting to see if the
|
||||
app came up, the regular backup copied the broken version into the app's local backup. It now
|
||||
leaves an app alone while it is updating. Live as 0.238.1.
|
||||
|
||||
15. **The demo machines no longer get new releases by themselves between golden images. One decision.**
|
||||
This morning we agreed to bake the golden image weekly, on the promise that every release still
|
||||
reaches both machines in about 20 seconds. **That promise was wrong.** The hub will not move a
|
||||
machine to a release newer than the golden image, so between bakes I install releases on the demo
|
||||
machines by hand. Today I did that three times.
|
||||
- **Bake a golden image for every release again.** Releases reach the machines by themselves. The
|
||||
cost is the weekly routine we just ended.
|
||||
- **Let the hub move machines past the golden image when a release says it needs no newer agent.**
|
||||
Releases reach the machines by themselves and goldens stay weekly. It needs one small hub change,
|
||||
and first each release must reliably say which agent it needs (R-470).
|
||||
|
||||
**My pick: the second.** **If you do nothing:** nothing breaks; I keep installing by hand and
|
||||
saying so each time.
|
||||
|
||||
16. **Apps with no copy on a second drive cannot be updated. One decision.** The update's safety rule
|
||||
asks for the app's second-drive copy, as we decided on 2 September. On the HP, two apps (gokapi
|
||||
and nextcloud) have no such copy, so their Update button refuses. A box with only one drive would
|
||||
refuse every update.
|
||||
- **Keep it.** Only apps with a second-drive copy can be updated. The refusal tells the customer how
|
||||
to switch it on.
|
||||
- **Also accept the backup on the same drive.** Updates work on one-drive boxes. That backup is lost
|
||||
if that drive dies.
|
||||
|
||||
**My pick: keep it for now,** and revisit when the first one-drive customer exists. **If you do
|
||||
nothing:** those apps show the refusal and nothing else changes.
|
||||
|
||||
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
@@ -239,8 +287,10 @@ nothing.*
|
||||
- **GOLDENS ARE NOW WEEKLY AND BEFORE ANY INSTALL, NOT PER RELEASE — AND THE GATE KNOWS. DECIDED 2026-09-13.**
|
||||
In August I baked 25 goldens in 26 days, almost one per release, because the check trips on every
|
||||
release on purpose. From today: one golden a week, and always before a drill or a fresh install.
|
||||
Every release still reaches both demo machines in about 20 seconds — only the image a **new**
|
||||
machine starts from moves to a cadence. The check now reads a dated permission slip that runs out
|
||||
**Correction, same day:** I wrote that every release would still reach both demo machines in about
|
||||
20 seconds between bakes. **That is wrong.** The hub will not move the machines past the image a new
|
||||
machine starts from, so between bakes I install each release on the demo machines by hand. That is
|
||||
item 15 below, and it is yours to decide. The check now reads a dated permission slip that runs out
|
||||
after at most 14 days; while it is valid the check warns instead of refusing, and when it runs out
|
||||
the check is red again until someone bakes or renews. A dated slip cannot be forgotten — it just
|
||||
expires. Today's golden (0.236.0) is baked, checked three ways, and live.
|
||||
|
||||
Reference in New Issue
Block a user