R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2 skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question, why the data legs are deliberately not guarded, and why the capture job is not guarded either. 00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what the route can be relied on for. Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a DECISION and deliberately not acted on - should a documents-only push be subject to the golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus what happens if Viktor does nothing. The gate was NOT changed. R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is owed and is more urgent than the previous six. STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with its do-nothing outcome. Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was not installed in the guest and silently did nothing, and a session that expired mid-run so a POST did nothing).
This commit is contained in:
@@ -13,21 +13,33 @@ delivered; both demo machines are on it and nothing is waiting on you about this
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **Nothing about delivery — the golden train is current.** Golden **0.229.0** was baked, published,
|
||||
round-trip verified, vouched, and the fleet floor raised to 0.229.0 on 2026-08-31 (the second bake
|
||||
that day; 0.228.0 was the first). Both demo machines run it; **`demo-felhom` got there by itself**
|
||||
and re-registered its jobs without anyone touching it. A machine installed today receives 0.229.0
|
||||
and everything shipped today, **including the restore from the second drive**. Evidence:
|
||||
`documentation/tests/golden-0.229.0-2026-08-31/`.
|
||||
1. **A golden carrying 0.230.0 is owed.** Golden **0.229.0** is vouched and the fleet floor is 0.229.0
|
||||
— **and 0.229.0 is the build that deletes a good copy** (R-403, measured today). Both demo machines
|
||||
and any machine installed right now carry that defect. `demo-hp` has been updated to 0.230.0 by
|
||||
hand; `demo-felhom` has not. **If you do nothing:** the fix stays on one machine and a newly
|
||||
installed box gets the defect. Baking and vouching 0.230.0 and raising the floor closes it — the
|
||||
same three-field change as this morning.
|
||||
|
||||
2. **Nothing else about this release.** Everything in 0.229.0 is a fix to code that ships in the
|
||||
controller image; no customer action, no data migration, no credential change.
|
||||
|
||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
|
||||
skipped that check **six times**, each time for a written reason: it runs on every push to the
|
||||
website/documentation repository, including pushes that change nothing a machine installs.
|
||||
**A guard we correctly skip six times is teaching us to skip it.**
|
||||
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
|
||||
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
|
||||
**The case against:** the check was earned — a release went out while machines were still being
|
||||
installed with the previous one, three times in three days — and narrowing a guard is how the thing
|
||||
it was built for comes back.
|
||||
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
|
||||
I have NOT changed it; this is yours to decide and mine to build.
|
||||
|
||||
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
@@ -95,6 +107,20 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
||||
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
||||
off-site permanently.
|
||||
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
|
||||
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
|
||||
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
|
||||
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
|
||||
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
|
||||
folder exist* before copying over it, and an empty package is a folder that exists.
|
||||
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
|
||||
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
|
||||
machine, in the same state: **all seven files still there, byte for byte.**
|
||||
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
|
||||
it — it names the real date of the package instead. And after a restore from the second drive, the
|
||||
first drive's package is filled back in immediately, so the empty state that started all this cannot
|
||||
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
|
||||
is fenced.
|
||||
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
|
||||
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
|
||||
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
|
||||
@@ -151,16 +177,6 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
||||
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
||||
would be a customer whose upload is busy. The warning is there so that does not happen.
|
||||
- **After a restore from the second drive, the first drive's package is rewritten EMPTY** (R-403,
|
||||
found during the 0.229.0 drill, **not fixed**). Two seconds after the restore finished, the box's
|
||||
five-minute housekeeping rebuilt the first drive's package from a drive that had no data files on
|
||||
it, and wrote a package that lists nothing. The ordinary „Visszaállítás indítása" then read it and
|
||||
said, correctly and uselessly, that the backup held only settings. **What we did not test:** the
|
||||
nightly copy mirrors the first drive over the second one and deletes what is not there, so the
|
||||
next night could plausibly overwrite the good copy with the empty one. That is a reading of the
|
||||
code, not a measurement, and it is written down as unmeasured on purpose. **If you do nothing:**
|
||||
a customer who recovers from their second drive may find, the next morning, that the copy they
|
||||
recovered from has been replaced by an empty one. Settling it costs one test on a spare machine.
|
||||
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
||||
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
||||
is **under a year** away on the corrected measurement, not two.
|
||||
|
||||
Reference in New Issue
Block a user