db0812b6f2
gates / gates (push) Successful in 16s
Second full delivery of the day. golden_currency_gate.py went red -> green on the same command, so the --no-verify bypass declared on the previous push is now historical rather than standing. GOLDEN_VERSION 0.227.1 GOLDEN_SHA256 66754491dc9bd0130ef8ded9562f63c53a5ffdcfd91baa551141e55fa083ea32 size 657 403 203 B baked gitea.dooplex.hu/admin/felhom-controller:0.227.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP. The published bytes were downloaded back -- size and sha256 identical to what the bake reported -- and ./etc/felhom-controller- image was read OUT of the downloaded archive: felhom-controller:0.227.1. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Markers counted: docker OK (overlay2 = 1, mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. 404 pre-gate passed before the run and the script's own pre-delete agreed, so nothing was overwritten. Three-field vouch, all three checked: agent_version 0.130.0 >= min_agent 0.129.0 (NOT the R-216 shape), wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash. The R-120 gate on that POST passed on its own terms rather than being worked around. AND THE LINE WORTH KEEPING. demo-felhom self-updated 0.226.1 -> 0.227.1 in under 30 seconds and then logged: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST A box nobody deployed to now runs today's off-site integrity check on its own schedule. That is a floor DELIVERING rather than merely recording, observed instead of assumed -- and it is the strongest evidence R-242 has carried. Token hygiene: file->file, read inside the VM by a runner script, never on a command line (systemctl show ... grep -c -F token = 0). The leak grep on the committed log was PROVEN TO WORK before its 0 was believed. Teardown: guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out, VM powered off, disk reverted to virgin. R-242 now records the cadence as MEASURED: five convictions and two full bakes in one day. Every bypass declared, every debt paid -- and the pattern the row exists to name is exactly that a release and its delivery are separate acts. Its other half stays open: nothing gates the VOUCH itself.
157 lines
12 KiB
Markdown
157 lines
12 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-23 — you now hear about EVERY broken app, not just the first one each hour. The
|
|
hub deployed itself; nothing is waiting on you except the floor from the last release.**
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
|
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
|
> If it does not fit, it belongs in the register instead.
|
|
|
|
## Waiting on you
|
|
|
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
|
nothing.*
|
|
|
|
1. **How deep should the off-site check go?** (R-399). The box now checks its off-site store weekly.
|
|
The check it runs today reads the catalogue — it catches a missing or unreadable backup, and it does
|
|
**not** re-read the stored bytes, so it cannot see a file that has quietly rotted.
|
|
**What it costs to go deeper, measured on your own machine today, not guessed:**
|
|
the shallow check takes **35.0 s**; re-reading **all** the data takes **39.2 s**. Four seconds more.
|
|
That is because the time goes on talking to the off-site box, not on moving data — and it will stop
|
|
being true as the store grows, so this is worth re-measuring, not deciding once forever.
|
|
**If you do nothing:** the catalogue is checked weekly and the stored bytes are never re-read.
|
|
I can turn it on with one setting whenever you say.
|
|
|
|
2. **Nothing about delivery — the golden train is current.** Golden **0.227.1** was baked, published,
|
|
round-trip verified, vouched, and the fleet floor raised to 0.227.1 on 2026-08-30. Both demo
|
|
machines run it; **`demo-felhom` got there by itself** and started the new off-site check on its own
|
|
schedule without anyone touching it. A machine installed today receives 0.227.1 and everything
|
|
shipped today. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`.
|
|
|
|
3. **Nothing else.** Everything in the releases of 2026-08-30 is a fix to code that ships in the
|
|
controller image; no customer action, no data migration, no credential change.
|
|
|
|
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
|
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
|
as it misled one by an hour.
|
|
|
|
## Decided — and what would reopen each
|
|
|
|
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
|
|
asks. *(R-312)*
|
|
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
|
|
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
|
|
abandoned. *(R-313)*
|
|
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
|
|
outside a constructed test. *(R-303)*
|
|
|
|
## What works
|
|
|
|
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
|
|
**Which controller each box runs, and where the floor sits, is item 1 above and is not restated here** —
|
|
R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the
|
|
page then disagreed with itself about the thing an operator checks first. Ask the hub (`/hosts`,
|
|
`/configs`) or the box for what is live; a doc is never the authority on a version.
|
|
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
|
|
|
|
**The fleet, because two summaries have been misread:** five customer records, three machines.
|
|
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
|
|
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
|
|
**`tester-1` is a record with no machine.**
|
|
|
|
## Shipped
|
|
|
|
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
|
|
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
|
|
The whole-machine backups had their own checks; the copies holding your customers' documents and
|
|
photos had none, so we would have found a problem at restore time, with a customer waiting.
|
|
Now the box checks its own off-site store about once a week and tells you only if something is
|
|
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
|
|
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
|
|
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
|
|
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
|
|
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
|
|
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
|
|
a second check fired during the first, which correctly stepped aside without doing anything.
|
|
**Read item 2 under „Waiting on you" for what this check does NOT see.**
|
|
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
|
|
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
|
|
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
|
|
|
|
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
|
|
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
|
|
that it called the ordinary backup routine — **which always writes to the app's normal backup
|
|
filename first** — so the app's real backup was overwritten and then renamed away. Until the next
|
|
nightly run the app had **no database backup of its own**, and a local recovery in that window would
|
|
have told you the app never had a database. A comment in the code said this could not happen; it
|
|
could, and had been happening for four months. Proven fixed the only way it can be: the app's own
|
|
backup file is now **byte-identical before and after a restore**, on both database types.
|
|
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
|
|
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
|
|
already taken a copy of your live database — a good copy — and **nothing in the product could put it
|
|
back.** You were shown a filename. On one of the two database types it was worse: part of the
|
|
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
|
|
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
|
|
app is running. Proven on both database types, byte-identical both times.
|
|
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
|
|
app on a half-written database lets you type into it and makes the damage permanent — and you are
|
|
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
|
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
|
off-site permanently.
|
|
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
|
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
|
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
|
were never given a drive to choose. It was asking one question to answer two. Proven today on
|
|
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
|
|
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
|
|
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
|
|
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
|
|
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
|
|
back, because a restore that mentions only its file count is how a silent loss reads as a success.
|
|
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
|
|
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
|
|
app of 53 was affected**, established with a check first proved able to catch a planted second case.
|
|
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
|
|
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
|
|
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
|
|
healthy.
|
|
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
|
|
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
|
|
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
|
|
The off-site box is back to **17** open connections from **415**.
|
|
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
|
|
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
|
|
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
|
|
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
|
|
can be re-attached after a reinstall** (R-280).
|
|
|
|
## Broken, or knowingly incomplete
|
|
|
|
- **The off-site store is only checked SHALLOWLY, and the deep check is switched off** (R-399).
|
|
This is the honest version of what shipped today. The box now checks its own off-site store about
|
|
once a week (R-359, controller 0.227.0) — and the check it runs reads the *catalogue* of the backups,
|
|
not the backups themselves. **We proved the difference:** a copy was damaged in a way that left its
|
|
size unchanged, and the shallow check said „no errors were found". Only the deep check caught it.
|
|
The deep check is built and **off**, waiting for your decision — see item 2 under „Waiting on you".
|
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
|
is **under a year** away on the corrected measurement, not two.
|
|
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
|
|
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
|
|
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
|
|
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
|
|
resolution goes missing quietly.
|
|
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
|
|
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
|
|
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
|
|
fixed twice** (R-327).
|
|
|
|
## Working on next
|
|
|
|
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
|
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|