R-672 session: audit README (not done first, wrong claims), STATUS, capability map, register 344->341, topic REPORT
gates / gates (push) Successful in 28s
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,28 +1,25 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-24 (night shift, run in daytime from 11:07). The demo boxes are on controller 0.269.1 — except demo-hp's customer box, which is damaged and needs you (first item under "What needs you").**
|
||||
**Updated 2026-09-24 (evening). The demo boxes are on controller 0.270.0. The HP customer box is repaired. A fixed host agent (0.133.0) is released and waits for your signature. Until then the automatic restore test is off on both demo boxes.**
|
||||
|
||||
**Decisions I took on my own (you may reverse either).**
|
||||
- **A second controller release tonight (0.269.1).** The first one (0.269.0) had a flaw I found in its own live test: when the catalog re-tested an image, the box wrote the new image into a running app's file before anyone pressed Update, so any restart would have changed the image with no backup and no undo. I fixed it and released again rather than send the flaw to the fleet. Now an installed app keeps its image until an Update moves it.
|
||||
- **A crash loop is 6 restarts in 10 minutes, not the 10 you wrote.** Measured: Docker slows a steady crash loop to about one restart a minute, so 10 would never catch it. No healthy app in any test record restarted more than once while starting.
|
||||
**Decisions I took on my own (you may reverse them).**
|
||||
- **The space check uses the real, uncompressed size, not the backup file size.** The rule in your brief ("file size × 1.2 + 5 GB") would NOT have stopped last night's accident. The HP box's backup file is 7 GB, but restoring it writes 22.6 GB.
|
||||
- **To switch the test off I used the value −1, not 0.** In this setting, 0 means "every 6 hours". Only a negative number turns it off.
|
||||
- **The hub got a small release too (0.124.0).** A thin disk pool now counts as urgent at 90 % (data or its bookkeeping part), with one alarm per pool per 6 hours. Before, the alarm came only at 95 %, up to 15 minutes late, and one full pool could silence another.
|
||||
|
||||
**What I tested, and it worked.**
|
||||
- **The second drive now brings an app with files back whole** (your ruling). Tested four times on the scratch machine, twice under a power cut or a nearly full disk. Every time: the account and the files came back, a file the household had edited later was kept, and nothing was deleted.
|
||||
- **A stranded app can only be removed keeping its data** (your ruling). The box refuses to delete the data, in both languages.
|
||||
- **The box stops a crash loop or a memory storm** (your ruling), tells the household and you, and Start gives one more try. A second stop within a day says support is informed.
|
||||
- **Exact image fingerprints.** The box runs exactly the image the catalog tested, and says "update available" when a newer tested image exists for the same version name.
|
||||
- **Automatic updates: measured, not built.** A test caller updated three apps (up to three steps each) and set a failing one aside in about 7 minutes. The build plan is written with the numbers.
|
||||
- **A chaos hour, 12 rounds** (power cuts, killed controller, Docker restarts, a full disk, backups). Updates resumed or undid themselves correctly every time.
|
||||
**What I did, and it worked.**
|
||||
- **The HP customer box is repaired.** I stopped it, checked both disks (small damage fixed, second check clean), and started it. All apps came back, it took the current version by itself, and the hub shows it online. Two apps' cache files (Redis) had been cut off by the full disk; with your yes I kept copies and repaired them. Both apps run.
|
||||
- **The box's own crash-loop stop worked for real:** while the cache was broken, the box stopped those two apps and told the hub, by itself.
|
||||
- **The new agent never starts a restore test that does not fit.** Tested on the HP box: it refused, said why, and created nothing.
|
||||
- **Four controller fixes (0.270.0):** Update on an app that is already current is refused, with no restart. An install cut off by a restart is now cleaned up, reported and shown on the page. After a restore, the box keeps the right health check. One log line is corrected.
|
||||
|
||||
**What broke.**
|
||||
- **demo-hp's customer box (not caused by tonight's work).** This morning the host's automatic restore test copied a whole guest onto the same full disk pool. The pool filled, and the customer box's disks became read-only. I freed the pool with an agent restart (the product's own clean-up). I could not repair the box itself.
|
||||
- **An install interrupted by a restart is lost without a word**, and a remove interrupted the same way is left half-done. Both recorded, not fixed.
|
||||
- After a failed update is restored, the box can keep the failed version's health check, and a later undo then fails for no reason. Recorded.
|
||||
- Not done: moving more apps in the catalog. It needs a throwaway test machine, and this session could not delete one afterwards.
|
||||
**What broke, or is not done.**
|
||||
- **The HP box's own backups have been failing every day since yesterday.** Its host disk has 4 GB free, but one backup needs 7 GB, and three old ones are kept. Nothing fixes this by itself.
|
||||
- **No full restore test fits on the HP box today** (it needs 30 GB free, 22 GB is free). The new agent will refuse it every time, correctly. So a passing test on the HP box is not proven yet.
|
||||
|
||||
**Rows.** 16 opened, 7 closed. The list went from 335 to 344.
|
||||
**Rows.** 1 opened, 4 closed, 2 updated. The list went from 344 to 341.
|
||||
|
||||
**What needs you.**
|
||||
1. **Repair demo-hp's customer box:** stop it, check its two disks, start it. If you do nothing, it keeps running but cannot save anything: no backups, no logs, no updates, and it stays on the old controller.
|
||||
2. **Decide how the restore test may use disk space** (for example: it must check free space first and never use the pool of the box it tests). If you do nothing, the next scheduled restore test can fill the pool again, on any box with a small disk.
|
||||
3. **Optional:** the two decisions above. If you do nothing, they stay as built.
|
||||
1. **Sign the agent update (0.133.0) for demo-hp and demo-felhom.** Version 0.133.0, sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, one `felhom-opsign -op agent_update` job per box, the same way as 0.131.0. If you do nothing, the boxes keep the old agent and the automatic restore test stays off.
|
||||
2. **After the new agent is on a box, turn the restore test back on.** On that host, remove the key `restore_test_eval_interval_seconds` from `backup` in `/etc/felhom-agent/agent.json` (the saved copy `agent.json.pre-r672` has it absent), then `systemctl restart felhom-agent`. If you do nothing, no backup is proven restorable on that box.
|
||||
3. **Decide what to do with the HP box's backup storage:** keep fewer old backups, back up somewhere else, or add disk. If you do nothing, its whole-box backup keeps failing every night.
|
||||
|
||||
Reference in New Issue
Block a user