Files
felhom.eu/documentation/audits/undo-live-2026-09-23
admin 05ea21e918
gates / gates (push) Successful in 25s
The undo, built and proven live: controller v0.263.2 (09 decision 15)
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
  live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
  the backup, after it and seconds before the press read back; cut-off copy
  held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
  ruled; R-646 opened. STATUS asks the floor question.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 12:25:13 +02:00
..
…
…
…
…
…

The undo, live on 9202 — controller v0.263.0 → v0.263.2, 2026-09-23

Venue: scratch guest 9202 (demo-hp), drill catalog (admin/app-catalog-drill, reset to live cfcfe5278428 before and after), update.health_timeout: 90s. Method: endpoint-level — every act is the endpoint the UI invokes (deploy, backup, sync, rescan, update, start, remove, the language switch) and every page is fetched as HTML (/apps/<app>?lang=hu|en); no browser. The product performs the undo; nothing here does. Each edge is a real migrating image move with, in the DRILL template only, a health probe on a port the app does not answer. Seed A before the backup, seed B after it, seed C seconds before the press — C exists only in the undo's own last-second copy.

Not done, or changed

  • Three releases, not one. v0.263.0 was built, deployed to 9202 only, and its first two live proofs FAILED honestly (HELD, data put back, old version judged "did not start"). Two defects the unit tests could not see; each fixed, red-proofed and released the same session: v0.263.1 (the undo's probe was gated on running while the current probe held the app unhealthy) and v0.263.2 (the "old" .felhom.yml saved at update time was already the new one — it flows in on every catalog sync). Only 9202 ever ran 0.263.0/0.263.1; both images were removed from it by name. No floor was raised.
  • The mail and the operator event of decision 15 are not built (09 §6.4 part 2).
  • The hold sentence body stays Hungarian on an English box (R-606); only the new prefix follows the box language.

Results (all on v0.263.2 unless stated)

proof result evidence
docmost (PostgreSQL) undone by the product failed probe at +107 s → undoing → undone at +137.6 s; A, B, C read back; 42 tables, 48 ledger before and after; 0 copies left 40-undo-docmost.txt
romm (MariaDB) undone at +163.9 s (undo 52 s); A, B, C; 24 tables, alembic 0095 before and after 40-undo-romm.txt
vikunja (SQLite in a volume) undone at +97.1 s (undo 3 s); A, B, C; 36 tables, ledger 117 before and after 40-undo-vikunja.txt
the page line, both languages hu: „A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el." · en: "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost." (same for romm, vikunja) 40-undo-*.txt
a cut-off copy (vikunja: the finished-marker taken from one copy during verifying) undoing → failed in 0.6 s, nothing poured back; hold (hu box): „A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése …"; en box: "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja …"; both copies kept 60-cutoff-vikunja.txt
a power cut during the undo (romm: pct stop 9202 the moment the phase read undoing) after boot: update recovery: romm was interrupted while UNDOING … RESUMING the undo → undone; A, B, C; ledger equal; 0 copies left 50-powercut-romm.txt
a person presses Update after an undo (docmost; the drill catalog then fixed the probe) done at +36.8 s on 0.96.0; A, B, C; the undone line gone from both pages; last_update_undone gone from app.yaml 70-manual-press-docmost.txt
a failed undo holds honestly (v0.263.0, docmost and romm — the defect above) HOLD „… Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el." — true: the data had been put back; the old version was never probed 10-docmost-undo.txt, 20-romm-undo.txt
removal deletes kept copies (vikunja, held with 2 copies) 2 → 0 after POST /api/stacks/vikunja/remove 95-teardown.txt
R-642, the Start answer Stack romm start requested — state now: running 80-r642-start-answer.txt

Extra downtime of the copy (the app is stopped there anyway to be recreated): the copying phase lasted 2.1 s (docmost), 6.8 s (romm, incl. a 3.8 s stop), 1.6 s (vikunja).

Teardown — three layers

  • machine (9202): docmost, romm, vikunja removed through the product (no containers, no volumes, no undo copies); romm's drive folder removed by name (the product kept it, R-442 as in every drill); test images removed by name (docmost 0.95.0/0.96.0, romm 5.0.0/5.3.0, vikunja 2.3.0/2.6.0, nextcloud 34.0.1-apache, controller 0.263.0/0.263.1); controller.yaml restored from the saved copy and read back identical; catalog cache re-cloned from the live repo (cfcfe52). 9202 stays on controller 0.263.2 (self-update off; the fleet floor is untouched at 0.262.1). Apps afterwards: the same three as at the start of the day (gokapi still crash-looping, R-644).
  • host (demo-hp): nothing provisioned; the guest was stopped and started once for the power cut.
  • hub: nothing touched.
  • drill repo: reset to live main, image lines identical, has_actions: false.