2263245cf2
gates / gates (push) Successful in 17s
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e, 657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/. WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a good copy. Every fresh install and the whole fleet floor still carried it. golden_currency_gate.py had been red acrossdddcc80,6e550ae,130f7a6and32a4c35. THREE INDEPENDENT READERS agreed before anything was vouched: the bake's own print, the round trip of the PUBLISHED bytes (HTTP 200, 657873700 B, same sha), and the hub's Day-0 dropdown reading Gitea on a different code path. And the delivered artifact names the controller it will start - ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.230.0, with 19382 entries under var/lib/felhom/docker/. BOTH PRE-GATES were shown able to see something before their zeroes were believed: the 404 pre-gate, and the token-leak grep which returns 0 on the committed log and 1 on a seeded throwaway copy. The transient unit's own properties were grepped for the token too - 0, with the same seeded positive control returning 1. Acceptance markers counted on the COMMITTED log: 1/1/1/1 present, 0/0 absent, and the zeroes are believable because the same including-mount-point pattern returns two real lines on that file. THE VOUCH IS A THREE-FIELD CHANGE and only one field moved, which is stated rather than left to look careless: golden_version 0.229.0 -> 0.230.0; agent_version 0.130.0 and min_agent 0.129.0 UNCHANGED because v0.230.0's CHANGELOG header says MinAgent 0.129.0 and 0.129.0 <= 0.130.0, so this is not the R-216 shape. The 303 flash was not treated as proof - the page was re-read and golden_behind_fleet confirmed absent. THE FLOOR is a separate setting and was raised on the operator's explicit answer: min_controller_version 0.229.0 -> 0.230.0. THE POSITIVE OBSERVABLE, from the agent's own journal on demo-felhom, which was still running the defective 0.229.0: 16:21:30 controller-swap: image file written, restarting bootstrap target=...0.230.0 16:21:40 controller-swap: new controller healthy target=...0.230.0 Both boxes now 0.230.0 healthy. Honest note: the polling loop's first read already said 0.230.0, so the transition was not seen by the loop - the journal is the evidence. R-410 FILED, found while the gate went green: golden_currency_gate.py is satisfied by a DIRECTORY NAME (EVIDENCE_RE against os.listdir, :89,:123). I created the evidence directory before the bake finished and the gate would have passed at that moment. It already declares that it does not check the vouch; it does not declare that the bake check is a filename check. Fix: read the GOLDEN_SHA256= line out of the directory's bake.log, with a red-proof on an empty directory. R-242 updated - seventh debt, paid the same day, twice in one day. Teardown: pct destroy 9100 --purge, shred -u AFTER the log was copied out, poweroff, qemu confirmed exited with ps -eo comm (not pgrep -f, which self-matches), disk reverted to virgin. All 13 gates green - the first push this session that needed no --no-verify. Ceiling R-409 -> R-410.
215 lines
17 KiB
Markdown
215 lines
17 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-31 (third pass) — golden 0.230.0 is baked, vouched and delivered. Both demo
|
|
machines are now on the build that stops a good backup copy being deleted; `demo-felhom` moved
|
|
itself. Nothing is waiting on you about that release any more.**
|
|
|
|
**Earlier 2026-08-31 — I measured whether the box could test its own off-site
|
|
restore without you. It can, and it is cheap — but not in the shape we had written down, so
|
|
there is a decision for you in item 4. No product code changed.**
|
|
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
|
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
|
> If it does not fit, it belongs in the register instead.
|
|
|
|
## Waiting on you
|
|
|
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
|
nothing.*
|
|
|
|
1. **Nothing is waiting on you about the 0.230.0 release.** Golden **0.230.0** is baked,
|
|
published and vouched, and the fleet floor is raised to 0.230.0. **`demo-felhom` was still
|
|
running 0.229.0 — the build that deletes a good copy — and moved itself across, unattended, in
|
|
about two minutes.** Both machines are healthy on 0.230.0, and a machine installed from scratch
|
|
now gets the fix too. Reversible if it ever needs to be: re-select the old values and save.
|
|
|
|
2. **Nothing else about this release.** Everything in 0.230.0 is a fix to code that ships in the
|
|
controller image; no customer action, no data migration, no credential change.
|
|
|
|
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
|
|
skipped that check **seven times**, each time for a written reason: it runs on every push to the
|
|
website/documentation repository, including pushes that change nothing a machine installs.
|
|
**A guard we correctly skip seven times is teaching us to skip it.**
|
|
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
|
|
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
|
|
**The case against:** the check was earned — a release went out while machines were still being
|
|
installed with the previous one, three times in three days — and narrowing a guard is how the thing
|
|
it was built for comes back.
|
|
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
|
|
I have NOT changed it; this is yours to decide and mine to build.
|
|
|
|
4. **Whether to have the box check its own off-site RESTORE every night** (R-87). I measured it
|
|
today instead of guessing. It is cheap: restoring **every** app on `demo-hp` — 8 backups, 774 MB —
|
|
took **25 seconds**, less than the 40 seconds the weekly check beside it already takes. But it
|
|
would catch **one** of the five restore faults we found by hand in the last six days, so the
|
|
version the old note asked for is not worth building.
|
|
**The version that IS worth building is a different question:** the weekly check proves the stored
|
|
bytes are the stored bytes. It cannot tell us we stored the **wrong thing** — an empty recovery
|
|
package backs up, checks and restores perfectly and gives the customer nothing back. That is not a
|
|
theory; it happened on 31 August (R-403). A nightly check of one app against its own packing list
|
|
would catch it and needs nothing new built underneath.
|
|
**If you do nothing:** the weekly check keeps being right about the bytes, and the first empty
|
|
package will be found by a customer trying to restore.
|
|
**My pick:** build the narrow version. **Yours to decide**, and I changed no code today.
|
|
|
|
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
|
|
|
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
|
as it misled one by an hour.
|
|
|
|
## Decided — and what would reopen each
|
|
|
|
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
|
|
asks. *(R-312)*
|
|
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
|
|
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
|
|
abandoned. *(R-313)*
|
|
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
|
|
outside a constructed test. *(R-303)*
|
|
|
|
## What works
|
|
|
|
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
|
|
**Which controller each box runs, and where the floor sits, is item 1 above and is not restated here** —
|
|
R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the
|
|
page then disagreed with itself about the thing an operator checks first. Ask the hub (`/hosts`,
|
|
`/configs`) or the box for what is live; a doc is never the authority on a version.
|
|
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
|
|
|
|
**The fleet, because two summaries have been misread:** five customer records, three machines.
|
|
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
|
|
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
|
|
**`tester-1` is a record with no machine.**
|
|
|
|
## Shipped
|
|
|
|
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
|
|
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
|
|
The whole-machine backups had their own checks; the copies holding your customers' documents and
|
|
photos had none, so we would have found a problem at restore time, with a customer waiting.
|
|
Now the box checks its own off-site store about once a week and tells you only if something is
|
|
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
|
|
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
|
|
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
|
|
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
|
|
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
|
|
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
|
|
a second check fired during the first, which correctly stepped aside without doing anything.
|
|
**Read item 2 under „Waiting on you" for what this check does NOT see.**
|
|
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
|
|
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
|
|
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
|
|
|
|
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
|
|
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
|
|
that it called the ordinary backup routine — **which always writes to the app's normal backup
|
|
filename first** — so the app's real backup was overwritten and then renamed away. Until the next
|
|
nightly run the app had **no database backup of its own**, and a local recovery in that window would
|
|
have told you the app never had a database. A comment in the code said this could not happen; it
|
|
could, and had been happening for four months. Proven fixed the only way it can be: the app's own
|
|
backup file is now **byte-identical before and after a restore**, on both database types.
|
|
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
|
|
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
|
|
already taken a copy of your live database — a good copy — and **nothing in the product could put it
|
|
back.** You were shown a filename. On one of the two database types it was worse: part of the
|
|
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
|
|
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
|
|
app is running. Proven on both database types, byte-identical both times.
|
|
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
|
|
app on a half-written database lets you type into it and makes the damage permanent — and you are
|
|
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
|
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
|
off-site permanently.
|
|
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
|
|
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
|
|
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
|
|
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
|
|
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
|
|
folder exist* before copying over it, and an empty package is a folder that exists.
|
|
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
|
|
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
|
|
machine, in the same state: **all seven files still there, byte for byte.**
|
|
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
|
|
it — it names the real date of the package instead. And after a restore from the second drive, the
|
|
first drive's package is filled back in immediately, so the empty state that started all this cannot
|
|
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
|
|
is fenced.
|
|
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
|
|
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
|
|
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
|
|
drive — its settings, its database and its data. It did that for months. **Nothing could open those
|
|
copies.** No button, no screen, no command. That mattered most in the one fault the second drive
|
|
exists for: if the first drive dies, the package on it dies too, and the copy that survived could
|
|
not be read. For **45** of the 53 apps that is everything they own.
|
|
Now the same restore that always worked from the first drive can read the copy on the second one,
|
|
and the button is on the app's own backup row. **Proved with the first drive's package taken away:**
|
|
Docmost came back in 29 seconds — its database, all three data volumes, a file with a Hungarian
|
|
accented name back byte for byte, and the app then read its own rows with its own password. Done
|
|
again with the app's password file also taken away: the copy carried the passwords too (2 of 2).
|
|
The screen that used to say „press that other button on another page" now offers the action itself.
|
|
It says plainly that **this one overwrites** what is there — the gentle „Fájlok visszaállítása"
|
|
beside it still only adds back missing files — and it names the date of the copy, so nobody puts
|
|
last week over today by accident.
|
|
- **We counted the apps this affects, and settled it.** Two of our own notes disagreed — 43 or 45.
|
|
The answer is **45**, counted with the product's own rule against the live catalogue. The older
|
|
count missed **radarr and sonarr**.
|
|
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
|
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
|
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
|
were never given a drive to choose. It was asking one question to answer two. Proven today on
|
|
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
|
|
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
|
|
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
|
|
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
|
|
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
|
|
back, because a restore that mentions only its file count is how a silent loss reads as a success.
|
|
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
|
|
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
|
|
app of 53 was affected**, established with a check first proved able to catch a planted second case.
|
|
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
|
|
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
|
|
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
|
|
healthy.
|
|
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
|
|
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
|
|
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
|
|
The off-site box is back to **17** open connections from **415**.
|
|
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
|
|
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
|
|
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
|
|
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
|
|
can be re-attached after a reinstall** (R-280).
|
|
|
|
## Broken, or knowingly incomplete
|
|
|
|
- **We do not know what the deep check costs on a BIG store** (R-401). Since 0.228.0 the weekly check
|
|
re-reads **all** your stored data, not just the list of it. We had to: a copy was damaged in a way
|
|
that left its size unchanged, and the old shallow check said „no errors were found". Only the deep
|
|
check caught it. **The cost we measured was four seconds** — 35.0 s before, 39.2 s after — but that
|
|
was on a 134 MB store, and it will not stay four seconds. So the box now tells us: any check that
|
|
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
|
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
|
would be a customer whose upload is busy. The warning is there so that does not happen.
|
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
|
is **under a year** away on the corrected measurement, not two.
|
|
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
|
|
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
|
|
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
|
|
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
|
|
resolution goes missing quietly.
|
|
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
|
|
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
|
|
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
|
|
fixed twice** (R-327).
|
|
|
|
## Working on next
|
|
|
|
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
|
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|