Decisions 151-156 recorded; design build evidence (R-638 measured safe, R-528 stopped by its measurement, R-518/R-717/R-762 built); golden 0.301.0; 07/08/09 updated; R-892, R-893 filed
gates / gates (push) Successful in 2m38s
gates / gates (push) Successful in 2m38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -593,6 +593,23 @@ files are neither deleted nor hidden — their visibility is a separate recorded
|
||||
`CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every
|
||||
capture — such as bounding the undo copies — has to sit ABOVE that check, not after it.
|
||||
|
||||
**[FACT] 2026-10-06 — restoring over a newer schema (R-638, `09` §3 decision 154).** The loader replays a copy ON TOP of
|
||||
the live database, so it removes only what the copy knows about. **Measured safe on the two main paths** (scratch 9202,
|
||||
controller 0.299.0, `audits/design-build-2026-10-06/B/`): a copy taken at the OLD version, then an Update that migrated,
|
||||
then the household's restore of that copy — docmost 0.95.0 → 0.96.0 (42 → 48 tables) came back with exactly the copy's
|
||||
42 and none of the 6 new ones; romm 5.0.0 → 5.3.0 (27 → 39) came back with exactly 27 (views counted), `Imported DB dump`
|
||||
logged, the data read back, the old version pinned. They are safe because the unit puts the copy's OWN volumes back
|
||||
before the replay. **Two side paths fixed by ORDER (controller v0.301.0):** the no-manifest fallback `RestoreApp` now
|
||||
starts only the database services, replays, then starts the app (it started the whole stack at the CURRENT definition
|
||||
first, so a newer app could migrate the old data before the replay); and a unit restore whose volume leg failed no
|
||||
longer replays (the copy's dump over a database volume that was NOT put back). The loader is unchanged; no delete step
|
||||
was added. **Known limit, not fixed (R-893):** after a failed OFF-SITE replay, the rollback loads the pre-restore copy
|
||||
(the NEWER state) over the OLDER volume just put back (`offbox_reconstitute.go` rollback), so tables the newer version
|
||||
removed stay, and when the snapshot's older definition was written, the rollback branch does not put the newer one
|
||||
back — the older app starts on rolled-back data. An order change cannot fix it: the only undo is a logical dump and the
|
||||
volume it should land in was replaced. Option B (a loader that rebuilds instead of overlays) or a pre-restore volume
|
||||
copy would; both are larger than this ruling.
|
||||
|
||||
**[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.**
|
||||
Recorded here rather than only in a closed register row, because a decision that survives only inside
|
||||
a closed work item is a decision nobody will find.
|
||||
@@ -672,7 +689,19 @@ successes only. After an agent restart the success is read back from the tier's
|
||||
|
||||
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
|
||||
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
|
||||
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
|
||||
**One stop per tier (controller v0.301.0, R-518 option A, `09` §3 decision 156 — this REVERSES R-82's „one quiesce
|
||||
window for both due tiers, never two app outages for one night").** A window runs only the FIRST due tier and starts
|
||||
the apps again at that tier's `snapshotted`; the upload then finishes with the apps running. Another due tier waits
|
||||
for a later cycle, in its own short window. A tier that refuses to start (BUSY or an error) inside a window still lets
|
||||
the next tier try in the same window — no copy has been made yet, so it is still one copy per stop. **The button
|
||||
(„Mentés most") makes the LOCAL copy only**, with one short stop; the off-site copy follows at the next night run. If
|
||||
the local tier's storage is absent the button makes no copy and stops nothing (logged at ERROR) — the off-site tier is
|
||||
never its stand-in. **Measured, read-only, 2026-10-06:** demo-felhom's night off-site job started 06:21:08 and reached
|
||||
`snapshotted` at 06:21:10, its one app running again at 06:21:18 — the off-site part of a stop is seconds; the rest is
|
||||
the apps' own stop and start (demo-hp 2026-10-05: stop 21 s, start 46 s). **Not shown live:** a press under the new
|
||||
rule (the scratch guest has no agent connection; the demo boxes take only deliveries and read-backs) — the first
|
||||
night run on the demo boxes is the proof owed (R-518). The page states about 1–1.5 minutes (an estimate from the
|
||||
parts above, not a measured press).
|
||||
**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the
|
||||
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
|
||||
longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both
|
||||
|
||||
@@ -236,6 +236,14 @@ was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202
|
||||
reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:**
|
||||
the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose
|
||||
`OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms.
|
||||
**Measured again 2026-10-06 on scratch 9202 (Docker 29.8.2; R-528's false flags were on 29.8.0, 2026-09-15): the flag
|
||||
was TRUE in all four shapes tried** — a child process killed while the container kept running (`oom_kill` 0 → 3,
|
||||
`OOMKilled=true`, still running) and three runs where the kernel also killed the main process (exit 137,
|
||||
`OOMKilled=true`); `docker events` carried an `oom` event each time. So R-528's starting point did not reproduce on
|
||||
today's engine, and its build (read the counter of every container; „probably out of memory" on exit 137) was STOPPED
|
||||
before any code, by the brief's measure-first rule — `09` §3 decision 155 stands for the operator to confirm or drop.
|
||||
Cost measured for the record: one `docker exec … cat memory.events` across demo-hp's 21 containers takes 1.6 s wall
|
||||
(0.47 s user, 0.44 s sys); cloudflared has no `cat` and reads as unknown. `audits/design-build-2026-10-06/C/`.
|
||||
|
||||
**The box STOPS a crash loop or an out-of-memory storm (decision 28 of `09` §3, R-667, controller v0.269.0 /
|
||||
hub v0.123.0, 2026-09-24).** Alarming was not enough: gokapi crash-looped for hours at 385 → 546 restarts
|
||||
|
||||
@@ -565,6 +565,16 @@ R-636's louder repeated alarm.
|
||||
is case-insensitive (termix's router ignores case). Final trick run: 113 tries on 11 apps, 0 got in.
|
||||
Same day, operator: CC changes the admin passwords of demo-hp's installed bookstack and calibre-web and stores them
|
||||
in the operator's credentials file (not in any repo).
|
||||
**2026-10-06 (R-717, controller v0.301.0): the window reopens a switch a COMMAND closed.** `after_setup` gains
|
||||
`open_command` / `open_success` (same `service`, `user`, `args_env` and argv-safe rules as `command`). The window
|
||||
marks the lock `opening` on disk first, lifts the env and runs `open_command`; only full success records `lifted`.
|
||||
A failed open closes it again at once and the app page says so (`app_info.signup_native_open_failed`). The close
|
||||
(`command`) runs when the window ends, at every controller start (the loop's first pass) and after a successful
|
||||
update (`markNativeLockForReapply`); a failed close is retried every 2 min and logged at ERROR. A template with
|
||||
`command` and no `open_command` keeps today's window (address block + env) and logs once that its own switch cannot
|
||||
reopen. **Wishlist** uses it (node's own sqlite module, `system_config.enableSignup` in group `global`), proven live
|
||||
on 9202. **Opengist cannot**: its container has no sqlite tool and no script runtime, and its CLI has no settings
|
||||
command — it keeps the address block alone (R-717 narrowed).
|
||||
50. **Every newly installed app goes off-site by itself when the customer has off-site** — *operator ruling 2026-09-30
|
||||
(R-720, option A).* It restores the intent of the 2026-09-16 ruling (`07` §6: Tier 3 ON for every new customer),
|
||||
which a per-app switch starting OFF had undone. If the apps will not fit the customer's quota, the page says so,
|
||||
@@ -905,6 +915,27 @@ its length, and both fixes cost something the household would notice — operato
|
||||
127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer.
|
||||
*Operator ruling 2026-10-05.* (R-861)
|
||||
|
||||
### 2026-10-06 (14:24) — three operator rulings and the reviewer's three design picks (recorded before the work)
|
||||
|
||||
151. **R-469 — closed.** The engine-major gate stays as decision 35's permanent per-app check. *Operator ruling
|
||||
2026-10-06 14:24* (option A).
|
||||
152. **The agent repo gets its copy of the shared rule file** (`.claude/rules/unprompted-work.md`), identical to the other
|
||||
copies. It adds rules and loosens nothing. *Operator ruling 2026-10-06 14:24* (option A).
|
||||
153. **The design items are built in the next session** (option A of the reviewer's 2026-10-06 proposal). *Operator
|
||||
ruling 2026-10-06 14:24.*
|
||||
154. **R-638 — option A: measure, then fix the order on the three side paths.** If the measurement shows the two main
|
||||
restore paths are safe, the row closes on A alone; option B (rebuild instead of overlay) stays a note in the row. The
|
||||
reverse-direction rollback (exposure 3) is fixed in the same session if the order fix covers it; if not, it becomes
|
||||
a known limit in `07` §6 and a row. *Reviewer's pick, 2026-10-06, standing (the operator saw it and did not change it).*
|
||||
155. **R-528 — options A then C** (read the kill counter of every running container; the crash-loop alarm says
|
||||
„probably out of memory" on exit 137). Option B (the host-side agent read) waits until A + C have run a week on the
|
||||
demo boxes. *Reviewer's pick, 2026-10-06, standing.*
|
||||
156. **R-518 — option A: one stop per backup tier.** Each window runs only the first due tier and resumes the apps at its
|
||||
`snapshotted`; the next due tier waits for a later cycle. A manual press makes the LOCAL copy only, with one short
|
||||
stop; the off-site copy follows at the next night run. **This reverses the recorded R-82 choice „ONE quiesce window
|
||||
for both due tiers (never two app outages for one night)".** *Reviewer's pick, 2026-10-06, standing — the operator's
|
||||
chat answer decides it; the brief carries it as A.*
|
||||
|
||||
### 2026-10-06 (13:25) — two operator rulings (recorded before the work)
|
||||
|
||||
149. **R-890 — the admin seed is allowed on the test boxes** (scratch 9202 and the disposable Tester 1 box) as well as
|
||||
|
||||
Reference in New Issue
Block a user