17d92e71a1
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run.
231 lines
18 KiB
Markdown
231 lines
18 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
|
|
|
**Updated 2026-09-01 (second pass) — both faults from the overnight test are fixed, and the golden
|
|
is baked, vouched and delivered. Both machines are on 0.232.0 and moved themselves. NOTHING is
|
|
waiting on you.**
|
|
|
|
**Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you,
|
|
then built the narrow version you picked. The measurement is why it is 3 seconds a night and
|
|
not an evening's work.**
|
|
|
|
|
|
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
|
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
|
> If it does not fit, it belongs in the register instead.
|
|
|
|
## Waiting on you
|
|
|
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
|
nothing.*
|
|
|
|
1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on
|
|
the real machines:
|
|
- the background job that could delete a live restore's lock now waits its turn — and the check
|
|
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
|
- the nightly backup check now runs on `demo-felhom`. It said **pass** there this morning, on a
|
|
machine where it could not start at all yesterday.
|
|
The golden carrying both is baked, vouched, and the floor is raised. `demo-felhom` picked it up
|
|
**by itself in about 20 seconds**.
|
|
|
|
2. **Whether to keep the test app `bentopdf` on `demo-hp`.** I deployed it last night because it
|
|
is the ONLY app of our 53 with neither a database nor stored files — which makes it the only
|
|
way to prove the new check does not cry wolf on an app that legitimately has nothing. It
|
|
passed silently, which is what we needed to see. **My pick: keep it**, as a permanent control.
|
|
**If you do nothing:** it stays, using almost no space. Say the word and I remove it.
|
|
|
|
3. **Nothing else about this release.** Everything in 0.232.0 ships in the controller image plus
|
|
two register lines in the hub (already live). No customer action, no data migration, no
|
|
credential change.
|
|
|
|
4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on
|
|
the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in
|
|
the control panel: seven of them, one a day. **And they are better than we thought** — I tried to
|
|
write into the snapshot area from a customer's machine and the storage **refused**, while the same
|
|
write to its normal folder worked. So a machine that wipes its own backup **cannot touch the
|
|
snapshots of it**. The worst case is losing about a day, then copying the rest back file by file.
|
|
**That is much smaller than what the notes have said since July.** I have corrected the notes.
|
|
**What I still need from you:** the machines can see the snapshot *door* but not what is inside —
|
|
only the main account can. So getting data back is you, in a browser, for now. **If you read one
|
|
snapshot's name off the panel and send it to me, one command settles whether the machines can
|
|
reach them directly** — and if they can, recovery becomes something the product does by itself.
|
|
**If you do nothing:** it stays a manual job for you, which is workable but slow.
|
|
|
|
5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a
|
|
test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it
|
|
once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`.
|
|
**No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one
|
|
mail.
|
|
|
|
6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
|
|
|
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
|
as it misled one by an hour.
|
|
|
|
## Decided — and what would reopen each
|
|
|
|
- **The missing-golden warning is now pointed at the people who can act on it (R-404). DECIDED
|
|
2026-09-01, and built the same day.** The warning was aimed at the wrong repository: the one where
|
|
a release actually happens never checked at all, while the one that only holds documents was
|
|
refused on every push — including the push that RECORDS a golden bake, which is the very act that
|
|
clears the warning. So the check was blocking its own cure, and we had skipped it thirteen times.
|
|
**What changed:** the release repository now prints a reminder the moment a release is committed
|
|
(it never blocks — you cannot bake a golden for a version you have not pushed yet), and the
|
|
documents repository still runs the check on every push and still says so loudly, but only refuses
|
|
a push that touches real code. **Every other check still blocks everything, always.** Nothing was
|
|
silenced and no product code changed. **Reopens if:** a release ever ships without a golden and
|
|
nobody noticed — that would mean the reminder is not reaching anyone, and the answer would be to
|
|
make the release repository refuse rather than remind.
|
|
|
|
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
|
|
asks. *(R-312)*
|
|
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
|
|
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
|
|
abandoned. *(R-313)*
|
|
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
|
|
outside a constructed test. *(R-303)*
|
|
|
|
## What works
|
|
|
|
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
|
|
**Which controller each box runs, and where the floor sits, is item 1 above and is not restated here** —
|
|
R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the
|
|
page then disagreed with itself about the thing an operator checks first. Ask the hub (`/hosts`,
|
|
`/configs`) or the box for what is live; a doc is never the authority on a version.
|
|
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
|
|
|
|
**The fleet, because two summaries have been misread:** five customer records, three machines.
|
|
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
|
|
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
|
|
**`tester-1` is a record with no machine.**
|
|
|
|
## Shipped
|
|
|
|
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
|
|
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
|
|
The whole-machine backups had their own checks; the copies holding your customers' documents and
|
|
photos had none, so we would have found a problem at restore time, with a customer waiting.
|
|
Now the box checks its own off-site store about once a week and tells you only if something is
|
|
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
|
|
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
|
|
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
|
|
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
|
|
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
|
|
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
|
|
a second check fired during the first, which correctly stepped aside without doing anything.
|
|
**Read item 2 under „Waiting on you" for what this check does NOT see.**
|
|
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
|
|
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
|
|
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
|
|
|
|
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
|
|
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
|
|
that it called the ordinary backup routine — **which always writes to the app's normal backup
|
|
filename first** — so the app's real backup was overwritten and then renamed away. Until the next
|
|
nightly run the app had **no database backup of its own**, and a local recovery in that window would
|
|
have told you the app never had a database. A comment in the code said this could not happen; it
|
|
could, and had been happening for four months. Proven fixed the only way it can be: the app's own
|
|
backup file is now **byte-identical before and after a restore**, on both database types.
|
|
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
|
|
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
|
|
already taken a copy of your live database — a good copy — and **nothing in the product could put it
|
|
back.** You were shown a filename. On one of the two database types it was worse: part of the
|
|
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
|
|
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
|
|
app is running. Proven on both database types, byte-identical both times.
|
|
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
|
|
app on a half-written database lets you type into it and makes the damage permanent — and you are
|
|
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
|
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
|
off-site permanently.
|
|
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
|
|
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
|
|
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
|
|
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
|
|
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
|
|
folder exist* before copying over it, and an empty package is a folder that exists.
|
|
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
|
|
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
|
|
machine, in the same state: **all seven files still there, byte for byte.**
|
|
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
|
|
it — it names the real date of the package instead. And after a restore from the second drive, the
|
|
first drive's package is filled back in immediately, so the empty state that started all this cannot
|
|
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
|
|
is fenced.
|
|
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
|
|
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
|
|
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
|
|
drive — its settings, its database and its data. It did that for months. **Nothing could open those
|
|
copies.** No button, no screen, no command. That mattered most in the one fault the second drive
|
|
exists for: if the first drive dies, the package on it dies too, and the copy that survived could
|
|
not be read. For **45** of the 53 apps that is everything they own.
|
|
Now the same restore that always worked from the first drive can read the copy on the second one,
|
|
and the button is on the app's own backup row. **Proved with the first drive's package taken away:**
|
|
Docmost came back in 29 seconds — its database, all three data volumes, a file with a Hungarian
|
|
accented name back byte for byte, and the app then read its own rows with its own password. Done
|
|
again with the app's password file also taken away: the copy carried the passwords too (2 of 2).
|
|
The screen that used to say „press that other button on another page" now offers the action itself.
|
|
It says plainly that **this one overwrites** what is there — the gentle „Fájlok visszaállítása"
|
|
beside it still only adds back missing files — and it names the date of the copy, so nobody puts
|
|
last week over today by accident.
|
|
- **We counted the apps this affects, and settled it.** Two of our own notes disagreed — 43 or 45.
|
|
The answer is **45**, counted with the product's own rule against the live catalogue. The older
|
|
count missed **radarr and sonarr**.
|
|
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
|
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
|
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
|
were never given a drive to choose. It was asking one question to answer two. Proven today on
|
|
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
|
|
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
|
|
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
|
|
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
|
|
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
|
|
back, because a restore that mentions only its file count is how a silent loss reads as a success.
|
|
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
|
|
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
|
|
app of 53 was affected**, established with a check first proved able to catch a planted second case.
|
|
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
|
|
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
|
|
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
|
|
healthy.
|
|
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
|
|
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
|
|
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
|
|
The off-site box is back to **17** open connections from **415**.
|
|
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
|
|
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
|
|
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
|
|
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
|
|
can be re-attached after a reinstall** (R-280).
|
|
|
|
## Broken, or knowingly incomplete
|
|
|
|
- **We do not know what the deep check costs on a BIG store** (R-401). Since 0.228.0 the weekly check
|
|
re-reads **all** your stored data, not just the list of it. We had to: a copy was damaged in a way
|
|
that left its size unchanged, and the old shallow check said „no errors were found". Only the deep
|
|
check caught it. **The cost we measured was four seconds** — 35.0 s before, 39.2 s after — but that
|
|
was on a 134 MB store, and it will not stay four seconds. So the box now tells us: any check that
|
|
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
|
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
|
would be a customer whose upload is busy. The warning is there so that does not happen.
|
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
|
is **under a year** away on the corrected measurement, not two.
|
|
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
|
|
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
|
|
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
|
|
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
|
|
resolution goes missing quietly.
|
|
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
|
|
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
|
|
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
|
|
fixed twice** (R-327).
|
|
|
|
## Working on next
|
|
|
|
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
|
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|