R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s

07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild
rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2
skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the
shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question,
why the data legs are deliberately not guarded, and why the capture job is not guarded either.

00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left
ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what
the route can be relied on for.

Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a
DECISION and deliberately not acted on - should a documents-only push be subject to the
golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus
what happens if Viktor does nothing. The gate was NOT changed.

R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect
in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses
git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is
owed and is more urgent than the previous six.

STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now
says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with
its do-nothing outcome.

Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was
not installed in the guest and silently did nothing, and a session that expired mid-run so a POST
did nothing).
This commit is contained in:
2026-08-31 14:39:29 +02:00
parent 66156c619f
commit dddcc808be
10 changed files with 195 additions and 22 deletions
+34 -18
View File
@@ -13,21 +13,33 @@ delivered; both demo machines are on it and nothing is waiting on you about this
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Nothing about delivery — the golden train is current.** Golden **0.229.0** was baked, published,
round-trip verified, vouched, and the fleet floor raised to 0.229.0 on 2026-08-31 (the second bake
that day; 0.228.0 was the first). Both demo machines run it; **`demo-felhom` got there by itself**
and re-registered its jobs without anyone touching it. A machine installed today receives 0.229.0
and everything shipped today, **including the restore from the second drive**. Evidence:
`documentation/tests/golden-0.229.0-2026-08-31/`.
1. **A golden carrying 0.230.0 is owed.** Golden **0.229.0** is vouched and the fleet floor is 0.229.0
— **and 0.229.0 is the build that deletes a good copy** (R-403, measured today). Both demo machines
and any machine installed right now carry that defect. `demo-hp` has been updated to 0.230.0 by
hand; `demo-felhom` has not. **If you do nothing:** the fix stays on one machine and a newly
installed box gets the defect. Baking and vouching 0.230.0 and raising the floor closes it — the
same three-field change as this morning.
2. **Nothing else about this release.** Everything in 0.229.0 is a fix to code that ships in the
controller image; no customer action, no data migration, no credential change.
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
skipped that check **six times**, each time for a written reason: it runs on every push to the
website/documentation repository, including pushes that change nothing a machine installs.
**A guard we correctly skip six times is teaching us to skip it.**
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
**The case against:** the check was earned — a release went out while machines were still being
installed with the previous one, three times in three days — and narrowing a guard is how the thing
it was built for comes back.
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
I have NOT changed it; this is yours to decide and mine to build.
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.
@@ -95,6 +107,20 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
database), and the undo copies no longer pile up forever — three per app, and they were being copied
off-site permanently.
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
folder exist* before copying over it, and an empty package is a folder that exists.
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
machine, in the same state: **all seven files still there, byte for byte.**
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
it — it names the real date of the package instead. And after a restore from the second drive, the
first drive's package is filled back in immediately, so the empty state that started all this cannot
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
is fenced.
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
@@ -151,16 +177,6 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
machine re-reads its whole store every week, however large it grows, and the first person to notice
would be a customer whose upload is busy. The warning is there so that does not happen.
- **After a restore from the second drive, the first drive's package is rewritten EMPTY** (R-403,
found during the 0.229.0 drill, **not fixed**). Two seconds after the restore finished, the box's
five-minute housekeeping rebuilt the first drive's package from a drive that had no data files on
it, and wrote a package that lists nothing. The ordinary „Visszaállítás indítása" then read it and
said, correctly and uselessly, that the backup held only settings. **What we did not test:** the
nightly copy mirrors the first drive over the second one and deletes what is not there, so the
next night could plausibly overwrite the good copy with the empty one. That is a reading of the
code, not a measurement, and it is written down as unmeasured on purpose. **If you do nothing:**
a customer who recovers from their second drive may find, the next morning, that the copy they
recovered from has been replaced by an empty one. Settling it costs one test on a spare machine.
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
is **under a year** away on the corrected measurement, not two.