Files
felhom.eu/STATUS.md
T
admin 0476a8d8e6
gates / gates (push) Successful in 17s
SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any
live store. ep0, DooPlex and Peti's box were not touched at all.

Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net
might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive
and a negative control on each: NO .snapshots is visible to either sub-account - not in the account
home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account
cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box
cannot restore from, so recovery would be an operator act at the Hetzner panel.

The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite
it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly
this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a
Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten
minutes in the panel, and it re-ranks everything.

Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a
progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388
AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly.

Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the
only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can
refuse; a Storage Box is a filesystem that runs nothing.

Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT
wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone
feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the
lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430.

Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag,
not a restic one.

Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS
reports, so the history to compare against is already on disk.

RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off
the box. Defer the transport change until Q1 is answered.

Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is
deliberately UNCHANGED.
2026-09-01 13:55:35 +02:00

225 lines
17 KiB
Markdown

# STATUS — what works, what's broken, what's next
**Updated 2026-09-01 (second pass) — both faults from the overnight test are fixed, and the golden
is baked, vouched and delivered. Both machines are on 0.232.0 and moved themselves. NOTHING is
waiting on you.**
**Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you,
then built the narrow version you picked. The measurement is why it is 3 seconds a night and
not an evening's work.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
> If it does not fit, it belongs in the register instead.
## Waiting on you
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on
the real machines:
- the background job that could delete a live restore's lock now waits its turn — and the check
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
- the nightly backup check now runs on `demo-felhom`. It said **pass** there this morning, on a
machine where it could not start at all yesterday.
The golden carrying both is baked, vouched, and the floor is raised. `demo-felhom` picked it up
**by itself in about 20 seconds**.
2. **Whether to keep the test app `bentopdf` on `demo-hp`.** I deployed it last night because it
is the ONLY app of our 53 with neither a database nor stored files — which makes it the only
way to prove the new check does not cry wolf on an app that legitimately has nothing. It
passed silently, which is what we needed to see. **My pick: keep it**, as a permanent control.
**If you do nothing:** it stays, using almost no space. Say the word and I remove it.
3. **Nothing else about this release.** Everything in 0.232.0 ships in the controller image plus
two register lines in the hub (already live). No customer action, no data migration, no
credential change.
4. **The copy that holds the customers' documents and photos can still be deleted by the box that
made it** (R-95 — first on the list since July, and this is the first time it has reached this
page). I studied it today and did not change anything. **One thing is yours and it takes ten
minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every
night and keeps seven. **I could not find a single one.** I looked from both machines, using the
credentials they already have, and checked that my method could see other things and could
correctly fail to see a made-up name. Either the snapshots are not being taken, or they are
invisible to the machines — and if they are invisible, they are also useless to them: getting one
back would be you, in the provider's control panel. **I did not log in to check, because your own
notes say that question is yours.** **If you do nothing:** the register keeps saying the net is
armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether
snapshots exist.** The answer changes which fix is worth building.
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.
## Decided — and what would reopen each
- **The missing-golden warning is now pointed at the people who can act on it (R-404). DECIDED
2026-09-01, and built the same day.** The warning was aimed at the wrong repository: the one where
a release actually happens never checked at all, while the one that only holds documents was
refused on every push — including the push that RECORDS a golden bake, which is the very act that
clears the warning. So the check was blocking its own cure, and we had skipped it thirteen times.
**What changed:** the release repository now prints a reminder the moment a release is committed
(it never blocks — you cannot bake a golden for a version you have not pushed yet), and the
documents repository still runs the check on every push and still says so loudly, but only refuses
a push that touches real code. **Every other check still blocks everything, always.** Nothing was
silenced and no product code changed. **Reopens if:** a release ever ships without a golden and
nobody noticed — that would mean the reminder is not reaching anyone, and the answer would be to
make the release repository refuse rather than remind.
- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer
asks. *(R-312)*
- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence
where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is
abandoned. *(R-313)*
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed
outside a constructed test. *(R-303)*
## What works
Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both.
**Which controller each box runs, and where the floor sits, is item 1 above and is not restated here** —
R-395: this paragraph carried a second copy of those numbers, it went stale by seven releases, and the
page then disagreed with itself about the thing an operator checks first. Ask the hub (`/hosts`,
`/configs`) or the box for what is live; a doc is never the authority on a version.
Off-site is credentialed on `demo-hp` and its store opens with the machine's own key.
**The fleet, because two summaries have been misread:** five customer records, three machines.
`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and
off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record.
**`tester-1` is a record with no machine.**
## Shipped
- **Something finally checks that the off-site copies are still there and readable** (R-359 + R-397,
controller 0.227.1, proven on `demo-hp`). Until today **nothing did** — not the box, not the agent.
The whole-machine backups had their own checks; the copies holding your customers' documents and
photos had none, so we would have found a problem at restore time, with a customer waiting.
Now the box checks its own off-site store about once a week and tells you only if something is
wrong. **A pass sends no e-mail, on purpose** — a weekly "everything is fine" is how people stop
reading their alerts. **It also catches itself up:** it asks „has it been more than seven days?",
not „is it Sunday?", so a machine that was switched off on its check day is checked the next day.
**And it never gets in the backup's way** — if a backup or restore is running, the check steps aside
and tries again tomorrow. That is not politeness: the tool would otherwise clear a lock that a live
backup was holding. Proven live, twice over — a real check against your real store (35 seconds), and
a second check fired during the first, which correctly stepped aside without doing anything.
**Read item 2 under „Waiting on you" for what this check does NOT see.**
- **The product stopped claiming a check it never ran.** The monitoring page said an integrity check
ran every Sunday. It did not exist. The e-mail text, the settings checkbox, the hub's side of it and
a debug button were all built and wired to nothing (R-397). They now have the missing piece.
- **Taking the safety copy no longer destroys the app's own backup** (R-361, controller 0.221.1,
proven on `demo-hp`). Before every restore the machine saves a copy of your live database. To do
that it called the ordinary backup routine — **which always writes to the app's normal backup
filename first** — so the app's real backup was overwritten and then renamed away. Until the next
nightly run the app had **no database backup of its own**, and a local recovery in that window would
have told you the app never had a database. A comment in the code said this could not happen; it
could, and had been happening for four months. Proven fixed the only way it can be: the app's own
backup file is now **byte-identical before and after a restore**, on both database types.
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
already taken a copy of your live database — a good copy — and **nothing in the product could put it
back.** You were shown a filename. On one of the two database types it was worse: part of the
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
app is running. Proven on both database types, byte-identical both times.
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
app on a half-written database lets you type into it and makes the damage permanent — and you are
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
database), and the undo copies no longer pile up forever — three per app, and they were being copied
off-site permanently.
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
folder exist* before copying over it, and an empty package is a folder that exists.
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
machine, in the same state: **all seven files still there, byte for byte.**
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
it — it names the real date of the package instead. And after a restore from the second drive, the
first drive's package is filled back in immediately, so the empty state that started all this cannot
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
is fenced.
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
drive — its settings, its database and its data. It did that for months. **Nothing could open those
copies.** No button, no screen, no command. That mattered most in the one fault the second drive
exists for: if the first drive dies, the package on it dies too, and the copy that survived could
not be read. For **45** of the 53 apps that is everything they own.
Now the same restore that always worked from the first drive can read the copy on the second one,
and the button is on the app's own backup row. **Proved with the first drive's package taken away:**
Docmost came back in 29 seconds — its database, all three data volumes, a file with a Hungarian
accented name back byte for byte, and the app then read its own rows with its own password. Done
again with the app's password file also taken away: the copy carried the passwords too (2 of 2).
The screen that used to say „press that other button on another page" now offers the action itself.
It says plainly that **this one overwrites** what is there — the gentle „Fájlok visszaállítása"
beside it still only adds back missing files — and it names the date of the copy, so nobody puts
last week over today by accident.
- **We counted the apps this affects, and settled it.** Two of our own notes disagreed — 43 or 45.
The answer is **45**, counted with the product's own rule against the live catalogue. The older
count missed **radarr and sonarr**.
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
were never given a drive to choose. It was asking one question to answer two. Proven today on
`privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files
back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet
visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way.
- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say
„0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came
back, because a restore that mentions only its file count is how a silent loss reads as a success.
- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355,
controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One
app of 53 was affected**, established with a check first proved able to catch a planted second case.
- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes,
hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it
would *not* have caught the 18 August fault, where one service was wedged and the machine stayed
healthy.
- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site
box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one
machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one.
The off-site box is back to **17** open connections from **415**.
- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on
the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per
secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the
installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive
can be re-attached after a reinstall** (R-280).
## Broken, or knowingly incomplete
- **We do not know what the deep check costs on a BIG store** (R-401). Since 0.228.0 the weekly check
re-reads **all** your stored data, not just the list of it. We had to: a copy was damaged in a way
that left its size unchanged, and the old shallow check said „no errors were found". Only the deep
check caught it. **The cost we measured was four seconds** — 35.0 s before, 39.2 s after — but that
was on a 134 MB store, and it will not stay four seconds. So the box now tells us: any check that
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
machine re-reads its whole store every week, however large it grows, and the first person to notice
would be a customer whose upload is busy. The warning is there so that does not happen.
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
is **under a year** away on the corrected measurement, not two.
- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent
since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is
lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled.
- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name
resolution goes missing quietly.
- **Three facts the machines send still have no reader** (R-264); **the storage page has its own
reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326:
23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we
fixed twice** (R-327).
## Working on next
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).