Files
felhom.eu/STATUS.md
T
admin 99af997ab9
gates / gates (push) Failing after 18s
R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A
pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth
that ships ON -- returned `no errors were found`, exit 0. Only --read-data
caught it. So the check that shipped verifies the index, the pack inventory and
the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a
bandwidth-and-cadence question; it is more than that, and its row now says so.

R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B /
2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none, because the wall clock is SFTP round-trips not transfer. The row
states the limit too: these do NOT extrapolate.

R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24
endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact
match, default NotFound -- so they 404. A third of a debug page does nothing, on
the surface an operator reaches for when something is already wrong.

R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying
resticStep is not a seam so no test can drive a restic path. The layer below it
has been injectable since the off-site tier shipped. The row survives as the
record that the seam EXISTS so nobody re-files it.

07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS
NOT R-359" because the two rows are adjacent and a check is not a restore-test.
08 alarm ladder: both event types recorded, including that `ok` is `info` and
therefore mails nobody BY DESIGN, and that all three registers were checked and
deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the
notifier and the hazard control; the scheduled firing is IMPLEMENTED only,
because a week has not passed.

wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The
gate was right -- the controller emits a field no hub struct can decode.
Building the display is a hub change and R-331 ruled that class the operator's
decision; the entry says to delete it when a surface exists.

This push used `git push --no-verify`. golden-currency is CONVICTED and right:
0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and
the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It
is item 3 under "Waiting on you".

Register 163 -> 165 -> 163.
2026-08-30 21:29:37 +02:00

162 lines
12 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-08-23 — you now hear about EVERY broken app, not just the first one each hour. The
hub deployed itself; nothing is waiting on you except the floor from the last release.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
> If it does not fit, it belongs in the register instead.
## Waiting on you
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Nothing — the golden train is current again.** Golden **0.226.1** was baked, published,
round-trip verified, vouched, and the fleet floor raised to 0.226.1 on 2026-08-30. Both demo
machines run 0.226.1; `demo-felhom` reached it by self-update, not by hand. A machine installed
today receives 0.226.1 and every fix from the four releases of 2026-08-30.
Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`.
2. **How deep should the off-site check go?** (R-399). The box now checks its off-site store weekly.
The check it runs today reads the catalogue — it catches a missing or unreadable backup, and it does
**not** re-read the stored bytes, so it cannot see a file that has quietly rotted.
**What it costs to go deeper, measured on your own machine today, not guessed:**
the shallow check takes **35.0 s**; re-reading **all** the data takes **39.2 s**. Four seconds more.
That is because the time goes on talking to the off-site box, not on moving data — and it will stop
being true as the store grows, so this is worth re-measuring, not deciding once forever.
**If you do nothing:** the catalogue is checked weekly and the stored bytes are never re-read.
I can turn it on with one setting whenever you say.
3. **Bake and vouch a golden carrying 0.227.1, then raise the floor** — the usual last step. `demo-hp`
runs 0.227.1; the fleet floor is 0.226.1 and the golden carries 0.226.1.
**If you do nothing:** a machine installed today gets 0.226.1 and none of today's off-site checking,
and `demo-felhom` stays where it is. Tracked on R-242.
4. **Nothing else.** Everything in the releases of 2026-08-30 is a fix to code that ships in the
controller image; no customer action, no data migration, no credential change.
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.
## Decided — and what would reopen each
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
asks. *(R-312)*
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
abandoned. *(R-313)*
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
outside a constructed test. *(R-303)*
## What works
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
**Which controller each box runs, and where the floor sits, is item 1 above and is not restated here** —
R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the
page then disagreed with itself about the thing an operator checks first. Ask the hub (`/hosts`,
`/configs`) or the box for what is live; a doc is never the authority on a version.
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
**The fleet, because two summaries have been misread:** five customer records, three machines.
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
**`tester-1` is a record with no machine.**
## Shipped
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
The whole-machine backups had their own checks; the copies holding your customers' documents and
photos had none, so we would have found a problem at restore time, with a customer waiting.
Now the box checks its own off-site store about once a week and tells you only if something is
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
a second check fired during the first, which correctly stepped aside without doing anything.
**Read item 2 under „Waiting on you" for what this check does NOT see.**
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
that it called the ordinary backup routine — **which always writes to the app's normal backup
filename first** — so the app's real backup was overwritten and then renamed away. Until the next
nightly run the app had **no database backup of its own**, and a local recovery in that window would
have told you the app never had a database. A comment in the code said this could not happen; it
could, and had been happening for four months. Proven fixed the only way it can be: the app's own
backup file is now **byte-identical before and after a restore**, on both database types.
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
already taken a copy of your live database — a good copy — and **nothing in the product could put it
back.** You were shown a filename. On one of the two database types it was worse: part of the
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
app is running. Proven on both database types, byte-identical both times.
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
app on a half-written database lets you type into it and makes the damage permanent — and you are
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
database), and the undo copies no longer pile up forever — three per app, and they were being copied
off-site permanently.
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
were never given a drive to choose. It was asking one question to answer two. Proven today on
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
back, because a restore that mentions only its file count is how a silent loss reads as a success.
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
app of 53 was affected**, established with a check first proved able to catch a planted second case.
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
healthy.
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
The off-site box is back to **17** open connections from **415**.
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
can be re-attached after a reinstall** (R-280).
## Broken, or knowingly incomplete
- **The off-site store is only checked SHALLOWLY, and the deep check is switched off** (R-399).
This is the honest version of what shipped today. The box now checks its own off-site store about
once a week (R-359, controller 0.227.0) — and the check it runs reads the *catalogue* of the backups,
not the backups themselves. **We proved the difference:** a copy was damaged in a way that left its
size unchanged, and the shallow check said „no errors were found". Only the deep check caught it.
The deep check is built and **off**, waiting for your decision — see item 2 under „Waiting on you".
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
is **under a year** away on the corrected measurement, not two.
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
resolution goes missing quietly.
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
fixed twice** (R-327).
## Working on next
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).