2f7c9a6ce5
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
132 lines
9.5 KiB
Markdown
132 lines
9.5 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-08-23 — the app-down alarm now actually reaches you by e-mail. It never has: 91 of
|
|
them were filed and not one was ever sent. Released and NOT yet delivered: step 1 is yours.**
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
|
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
|
> If it does not fit, it belongs in the register instead.
|
|
|
|
## Waiting on you
|
|
|
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
|
nothing.*
|
|
|
|
1. **Vouch the golden carrying controller 0.223.0** — Hub → Configuration → Day-0 artifacts.
|
|
**It is already baked, published and round-trip verified**
|
|
(`documentation/tests/golden-0.223.0-2026-08-23/`); only the vouch is left, and only you can do it.
|
|
**It is a THREE-field save:** `golden_version` → **0.223.0**, `agent_version` → **0.130.0**,
|
|
`min_agent` → **0.129.0**. **Then** raise the floor to **0.223.0**, last, in its own save.
|
|
**If you do nothing:** the fleet stays on 0.222.0, where the app-down alarm shows on the dashboard
|
|
and e-mails nobody, and where an app stopped from outside is still reported as if you had stopped
|
|
it yourself. New machines still receive 0.222.0.
|
|
*(The hub half is already live — v0.107.0 deployed itself through the manifest. Nothing owed there.)*
|
|
|
|
2. **A new switch has appeared for your customers, and it is OFF** — „Alkalmazás nem fut". **You get
|
|
the e-mail either way**; the switch only decides whether the customer also does. This is what you
|
|
asked for and it needs nothing from you. Mentioned so it is not a surprise the first time you see
|
|
the settings page.
|
|
|
|
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
|
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
|
as it misled one by an hour.
|
|
|
|
## Decided — and what would reopen each
|
|
|
|
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
|
|
asks. *(R-312)*
|
|
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
|
|
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
|
|
abandoned. *(R-313)*
|
|
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
|
|
outside a constructed test. *(R-303)*
|
|
|
|
## What works
|
|
|
|
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
|
|
`demo-hp` runs controller **0.219.0**; the fleet floor is still **0.218.0** (see item 2 above).
|
|
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
|
|
|
|
**The fleet, because two summaries have been misread:** five customer records, three machines.
|
|
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
|
|
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
|
|
**`tester-1` is a record with no machine.**
|
|
|
|
## Shipped
|
|
|
|
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
|
|
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
|
|
that it called the ordinary backup routine — **which always writes to the app's normal backup
|
|
filename first** — so the app's real backup was overwritten and then renamed away. Until the next
|
|
nightly run the app had **no database backup of its own**, and a local recovery in that window would
|
|
have told you the app never had a database. A comment in the code said this could not happen; it
|
|
could, and had been happening for four months. Proven fixed the only way it can be: the app's own
|
|
backup file is now **byte-identical before and after a restore**, on both database types.
|
|
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
|
|
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
|
|
already taken a copy of your live database — a good copy — and **nothing in the product could put it
|
|
back.** You were shown a filename. On one of the two database types it was worse: part of the
|
|
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
|
|
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
|
|
app is running. Proven on both database types, byte-identical both times.
|
|
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
|
|
app on a half-written database lets you type into it and makes the damage permanent — and you are
|
|
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
|
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
|
off-site permanently.
|
|
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
|
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
|
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
|
were never given a drive to choose. It was asking one question to answer two. Proven today on
|
|
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
|
|
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
|
|
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
|
|
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
|
|
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
|
|
back, because a restore that mentions only its file count is how a silent loss reads as a success.
|
|
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
|
|
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
|
|
app of 53 was affected**, established with a check first proved able to catch a planted second case.
|
|
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
|
|
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
|
|
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
|
|
healthy.
|
|
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
|
|
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
|
|
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
|
|
The off-site box is back to **17** open connections from **415**.
|
|
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
|
|
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
|
|
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
|
|
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
|
|
can be re-attached after a reinstall** (R-280).
|
|
|
|
## Broken, or knowingly incomplete
|
|
|
|
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not
|
|
the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the
|
|
standard tool — which we never run.
|
|
- **A restore that returns nothing still reports success** (part of R-354's neighbourhood, not fixed
|
|
today) and **verification copies have no delete guard**. Both deliberately left for their own rows.
|
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
|
is **under a year** away on the corrected measurement, not two.
|
|
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
|
|
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
|
|
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
|
|
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
|
|
resolution goes missing quietly.
|
|
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
|
|
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
|
|
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
|
|
fixed twice** (R-327).
|
|
|
|
## Working on next
|
|
|
|
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
|
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|