Files
felhom.eu/STATUS.md
T
admin 2263245cf2
gates / gates (push) Successful in 17s
golden 0.230.0 baked, vouched, floor raised - demo-felhom moved itself off the R-403 build (R-410 filed)
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e,
657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/.

WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a
good copy. Every fresh install and the whole fleet floor still carried it.
golden_currency_gate.py had been red across dddcc80, 6e550ae, 130f7a6 and 32a4c35.

THREE INDEPENDENT READERS agreed before anything was vouched: the bake's own print, the
round trip of the PUBLISHED bytes (HTTP 200, 657873700 B, same sha), and the hub's Day-0
dropdown reading Gitea on a different code path. And the delivered artifact names the
controller it will start - ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.230.0, with 19382 entries under var/lib/felhom/docker/.

BOTH PRE-GATES were shown able to see something before their zeroes were believed: the
404 pre-gate, and the token-leak grep which returns 0 on the committed log and 1 on a
seeded throwaway copy. The transient unit's own properties were grepped for the token
too - 0, with the same seeded positive control returning 1. Acceptance markers counted on
the COMMITTED log: 1/1/1/1 present, 0/0 absent, and the zeroes are believable because the
same including-mount-point pattern returns two real lines on that file.

THE VOUCH IS A THREE-FIELD CHANGE and only one field moved, which is stated rather than
left to look careless: golden_version 0.229.0 -> 0.230.0; agent_version 0.130.0 and
min_agent 0.129.0 UNCHANGED because v0.230.0's CHANGELOG header says MinAgent 0.129.0 and
0.129.0 <= 0.130.0, so this is not the R-216 shape. The 303 flash was not treated as
proof - the page was re-read and golden_behind_fleet confirmed absent.

THE FLOOR is a separate setting and was raised on the operator's explicit answer:
min_controller_version 0.229.0 -> 0.230.0. THE POSITIVE OBSERVABLE, from the agent's own
journal on demo-felhom, which was still running the defective 0.229.0:
  16:21:30 controller-swap: image file written, restarting bootstrap  target=...0.230.0
  16:21:40 controller-swap: new controller healthy                    target=...0.230.0
Both boxes now 0.230.0 healthy. Honest note: the polling loop's first read already said
0.230.0, so the transition was not seen by the loop - the journal is the evidence.

R-410 FILED, found while the gate went green: golden_currency_gate.py is satisfied by a
DIRECTORY NAME (EVIDENCE_RE against os.listdir, :89,:123). I created the evidence
directory before the bake finished and the gate would have passed at that moment. It
already declares that it does not check the vouch; it does not declare that the bake
check is a filename check. Fix: read the GOLDEN_SHA256= line out of the directory's
bake.log, with a red-proof on an empty directory.

R-242 updated - seventh debt, paid the same day, twice in one day.

Teardown: pct destroy 9100 --purge, shred -u AFTER the log was copied out, poweroff,
qemu confirmed exited with ps -eo comm (not pgrep -f, which self-matches), disk reverted
to virgin.

All 13 gates green - the first push this session that needed no --no-verify.

Ceiling R-409 -> R-410.
2026-08-31 16:25:35 +02:00

17 KiB

STATUS — what works, what's broken, what's next

Updated 2026-08-31 (third pass) — golden 0.230.0 is baked, vouched and delivered. Both demo machines are now on the build that stops a good backup copy being deleted; demo-felhom moved itself. Nothing is waiting on you about that release any more.

Earlier 2026-08-31 — I measured whether the box could test its own off-site restore without you. It can, and it is cheap — but not in the shape we had written down, so there is a decision for you in item 4. No product code changed.

A view, not a source. documentation/backlog/OPEN-ITEMS.md is the authority; this page restates part of it in plain words, and nothing may exist only here. Items, not paragraphs. One screen. If it does not fit, it belongs in the register instead.

Waiting on you

This section is allowed to be longer than one screen, and each item says what happens if you do nothing.

  1. Nothing is waiting on you about the 0.230.0 release. Golden 0.230.0 is baked, published and vouched, and the fleet floor is raised to 0.230.0. demo-felhom was still running 0.229.0 — the build that deletes a good copy — and moved itself across, unattended, in about two minutes. Both machines are healthy on 0.230.0, and a machine installed from scratch now gets the fix too. Reversible if it ever needs to be: re-select the old values and save.

  2. Nothing else about this release. Everything in 0.230.0 is a fix to code that ships in the controller image; no customer action, no data migration, no credential change.

  3. Whether a documents-only push should still be checked for a missing golden (R-404). We have now skipped that check seven times, each time for a written reason: it runs on every push to the website/documentation repository, including pushes that change nothing a machine installs. A guard we correctly skip seven times is teaching us to skip it. The case for narrowing it: a documents-only push cannot be the one that finishes a release, so only checking pushes that touch real code would fire on exactly the risky ones and end the habit. The case against: the check was earned — a release went out while machines were still being installed with the previous one, three times in three days — and narrowing a guard is how the thing it was built for comes back. If you do nothing: nothing breaks, the skipping stays routine, and the count keeps rising. I have NOT changed it; this is yours to decide and mine to build.

  4. Whether to have the box check its own off-site RESTORE every night (R-87). I measured it today instead of guessing. It is cheap: restoring every app on demo-hp — 8 backups, 774 MB — took 25 seconds, less than the 40 seconds the weekly check beside it already takes. But it would catch one of the five restore faults we found by hand in the last six days, so the version the old note asked for is not worth building. The version that IS worth building is a different question: the weekly check proves the stored bytes are the stored bytes. It cannot tell us we stored the wrong thing — an empty recovery package backs up, checks and restores perfectly and gives the customer nothing back. That is not a theory; it happened on 31 August (R-403). A nightly check of one app against its own packing list would catch it and needs nothing new built underneath. If you do nothing: the weekly check keeps being right about the bytes, and the first empty package will be found by a customer trying to restore. My pick: build the narrow version. Yours to decide, and I changed no code today.

  5. Whether to change the hub password (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. If you do nothing: it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.

  6. demo-hp's network setup does not match our own notes (R-338) — the machine works, the page is wrong, or the other way round. If you do nothing: the page keeps misleading the next session, as it misled one by an hour.

Decided — and what would reopen each

  • Getting old backups back yourself: NOT BUILT, deliberately. Reopens if: a real customer asks. (R-312)
  • The unopenable old copy on demo-felhom: KEPT as a test fixture — the only state in existence where a set-aside store is present and cannot be opened. Delete when: that work ships or is abandoned. (R-313)
  • A machine in two kinds of trouble says both things: LEFT AS IT IS. Reopens if: observed outside a constructed test. (R-303)

What works

Both demo machines are home, healthy and reporting — agent 0.130.0 published and running on both. Which controller each box runs, and where the floor sits, is item 1 above and is not restated here — R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the page then disagreed with itself about the thing an operator checks first. Ask the hub (/hosts, /configs) or the box for what is live; a doc is never the authority on a version. Off-site is credentialed on demo-hp and its store opens with the machine's own key.

The fleet, because two summaries have been misread: five customer records, three machines. demo-felhom and demo-hp are ours and disposable; drill-r50 is a nested drill VM, reverted and off. peti-felhom is a real machine we have not heard from since 15 July and has no host record. tester-1 is a record with no machine.

Shipped

  • Something finally checks that the off-site copies are still there and readable (R-359 + R-397, controller 0.227.1, proven on demo-hp). Until today nothing did — not the box, not the agent. The whole-machine backups had their own checks; the copies holding your customers' documents and photos had none, so we would have found a problem at restore time, with a customer waiting. Now the box checks its own off-site store about once a week and tells you only if something is wrong. A pass sends no e-mail, on purpose — a weekly "everything is fine" is how people stop reading their alerts. It also catches itself up: it asks „has it been more than seven days?", not „is it Sunday?", so a machine that was switched off on its check day is checked the next day. And it never gets in the backup's way — if a backup or restore is running, the check steps aside and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and a second check fired during the first, which correctly stepped aside without doing anything. Read item 2 under „Waiting on you" for what this check does NOT see.

  • The product stopped claiming a check it never ran. The monitoring page said an integrity check ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and a debug button were all built and wired to nothing (R-397). They now have the missing piece.

  • Taking the safety copy no longer destroys the app's own backup (R-361, controller 0.221.1, proven on demo-hp). Before every restore the machine saves a copy of your live database. To do that it called the ordinary backup routine — which always writes to the app's normal backup filename first — so the app's real backup was overwritten and then renamed away. Until the next nightly run the app had no database backup of its own, and a local recovery in that window would have told you the app never had a database. A comment in the code said this could not happen; it could, and had been happening for four months. Proven fixed the only way it can be: the app's own backup file is now byte-identical before and after a restore, on both database types.

  • A failed database restore now puts your data back by itself (R-379/R-380, controller 0.220.2, proven on demo-hp). Until today, if a restore of an app's database went wrong, the machine had already taken a copy of your live database — a good copy — and nothing in the product could put it back. You were shown a filename. On one of the two database types it was worse: part of the restore applied, part did not, and the dashboard said the app was healthy. Now the machine puts your own copy back automatically and says plainly: the restore failed, your data is as it was, the app is running. Proven on both database types, byte-identical both times. If even that fails, the app is deliberately stopped and held rather than started — a running app on a half-written database lets you type into it and makes the damage permanent — and you are told to contact us. That was your ruling this morning. Two things also stopped: the error no longer pastes raw database text at you (it was 615 bytes once, including rows out of your own database), and the undo copies no longer pile up forever — three per app, and they were being copied off-site permanently.

  • An empty package can no longer wipe out a good one (R-403, controller 0.230.0). Yesterday we wrote this down as suspected and said plainly it had not been tested. We tested it first, and it was real. On a demo machine, on yesterday's build: an app's copy on the second drive went from 120 MB — four database backups and three data archives — to 7 KB, nothing left, in a single nightly run, and the run reported success. The cause was that the nightly job only asked does the folder exist before copying over it, and an empty package is a folder that exists. Now the nightly job refuses to replace a complete package with an empty one. It keeps what it has, says so on the app's own backup page, and carries on with everything else. Proven on the same machine, in the same state: all seven files still there, byte for byte. Two more things came with it. The page no longer calls that copy fresh when the run did not refresh it — it names the real date of the package instead. And after a restore from the second drive, the first drive's package is filled back in immediately, so the empty state that started all this cannot happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case is fenced.

  • The copy on the second drive can now bring an app back (R-102 + R-103, controller 0.229.0, proven on demo-hp, and delivered — golden 0.229.0 is vouched and the fleet floor is raised, so a machine installed today has it). Every night the box copied each app's whole recovery package onto the second drive — its settings, its database and its data. It did that for months. Nothing could open those copies. No button, no screen, no command. That mattered most in the one fault the second drive exists for: if the first drive dies, the package on it dies too, and the copy that survived could not be read. For 45 of the 53 apps that is everything they own. Now the same restore that always worked from the first drive can read the copy on the second one, and the button is on the app's own backup row. Proved with the first drive's package taken away: Docmost came back in 29 seconds — its database, all three data volumes, a file with a Hungarian accented name back byte for byte, and the app then read its own rows with its own password. Done again with the app's password file also taken away: the copy carried the passwords too (2 of 2). The screen that used to say „press that other button on another page" now offers the action itself. It says plainly that this one overwrites what is there — the gentle „Fájlok visszaállítása" beside it still only adds back missing files — and it names the date of the copy, so nobody puts last week over today by accident.

  • We counted the apps this affects, and settled it. Two of our own notes disagreed — 43 or 45. The answer is 45, counted with the product's own rule against the live catalogue. The older count missed radarr and sonarr.

  • The off-site restore now works for the other 40 apps (R-356, controller 0.219.0, proven on demo-hp). It used to refuse before starting, tell the customer a running app „nincs telepítve", and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they were never given a drive to choose. It was asking one question to answer two. Proven today on privatebin: data planted through the app itself, backed up, deleted, restored — all 15 files back byte for byte, Hungarian accented names included, message „0 fájl és 1 adatkötet visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.

  • The off-site restore gives an app's data back at all (R-354, controller 0.218.0). It used to say „0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came back, because a restore that mentions only its file count is how a silent loss reads as a success.

  • Paperless's database is in the backup, and restoring it takes an undo copy first (R-355, controller 0.218.0). The dump was landing in a folder named after an app that does not exist. One app of 53 was affected, established with a check first proved able to catch a planted second case.

  • The system tells you when it cannot see the off-site copies (R-339) — a mail after ~30 minutes, hourly while it lasts, one all-clear. Caveat: it watches whether the machine answers, so it would not have caught the 18 August fault, where one service was wedged and the machine stayed healthy.

  • The connection leak was ours and is fixed (R-344). Our agent opened a connection to the off-site box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one. The off-site box is back to 17 open connections from 415.

  • A dated check can no longer be quietly missed (R-341) — but it speaks on the next push, not on the day. A machine we tell to be quiet is no longer reported as dead (R-321). One name per secret (R-295, R-323). The hub's own words are under a guard (R-324). Removal reverses the installation (R-316). A correct recovery code is no longer called wrong (R-311). The drive can be re-attached after a reinstall (R-280).

Broken, or knowingly incomplete

  • We do not know what the deep check costs on a BIG store (R-401). Since 0.228.0 the weekly check re-reads all your stored data, not just the list of it. We had to: a copy was damaged in a way that left its size unchanged, and the old shallow check said „no errors were found". Only the deep check caught it. The cost we measured was four seconds — 35.0 s before, 39.2 s after — but that was on a 134 MB store, and it will not stay four seconds. So the box now tells us: any check that takes longer than five minutes writes a warning naming this item. If you do nothing: every machine re-reads its whole store every week, however large it grows, and the first person to notice would be a customer whose upload is busy. The warning is there so that does not happen.
  • We ask the off-site box a question about once a second (R-336) — ~85,000 a day for a box we write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling is under a year away on the corrected measurement, not two.
  • Peti's machine has no recovery route at all. A real machine belonging to a real person, silent since 15 July, no key, no off-site copy, no local backup. If that drive fails, everything on it is lost. First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
  • The agent picks dnsmasq by looking at a file another package owns (R-317) — one line; LAN name resolution goes missing quietly.
  • Three facts the machines send still have no reader (R-264); the storage page has its own reason for an empty list (R-298); two thirds of the standing picture is unproven (R-326: 23 of 55 claims walked — python3 scripts/unproven.py); the picture still describes one defect we fixed twice (R-327).

Working on next

The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).