Files
felhom.eu/STATUS.md
T
admin 877fcd2a38
gates / gates (push) Successful in 16s
R-354 + R-355 CLOSED, proven live; golden 0.218.0 baked; R-367 filed
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.

R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.

R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.

Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.

R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.

Ceiling R-366 -> R-367.
2026-08-22 10:11:46 +02:00

219 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# STATUS — what works, what's broken, what's next
**Updated 2026-08-22 — both of last night's worst findings are fixed and proven on the machine.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
> If it does not fit, it belongs in the register instead.
## Waiting on you
*Nothing. All three questions that stood here were answered on 12–13 August and have moved to
**Decided** below. A decided question left in the deciding list is how a person loses track of what is
actually waiting.*
## Decided — and what would reopen each
*A decision with no trigger becomes a permanent silence, so each one names what would make us look
again.*
- **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a
support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks —
one request from a person who is not us. *(register: R-312)*
- **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the
only state in existence where a set-aside store is present and cannot be opened, which is the case
any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is
abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)*
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is
unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed
happening outside a constructed test. *(register: R-303)*
## What works
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.217.0, agent
0.130.0**. On `demo-hp` the floor delivered it unaided on 21 August: 0.216.0 → 0.217.0, **17 seconds**
from the operator pressing save to the new controller reporting healthy, with no customer action.
Off-site is credentialed on `demo-hp`, its repository opens with the machine's own key, and
`restic check` over the whole store reports **no errors**. `drill-r50` is reverted to `virgin`, powered
off. **`demo-hp` was reinstalled on 21 August and its off-site backup had been silent since 9 August**
— the rebuild lost the off-site target, and after that was healed every per-app off-site switch was
still off, so the nightly run reported "backup OK" having backed up nothing. Both are on again.
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom`
is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a
record with no machine** — created 13 August, no host, no backups, nothing to lose.
## Shipped
- **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a
completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks
only ever watched how full a store was getting, and a failed reading was written to a log nobody
reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the
weekly backup that happened to fall inside it. After about half an hour of being unable to see a
store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back.
**Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have
caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That
second check is written down as the next step.
- **The auto-update floor is current again** (R-343). You raised it to today's version this
afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*,
after the image is vouched, because it acts within seconds. It did: **one machine updated itself
nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not
paperwork.
- **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the
19th", that date is now read by the build system, and a push is refused once it passes. **It is not
a reminder service** — it speaks on the next push, not on the day — and that limit is written into
the check itself.
- **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image
had been two releases behind since the 14th — anyone installing would have received a version
missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked
and published, and **you vouched it**, which was the half that could not be done without you. **The
build system is green again for the first time since 14 August**, so the failure mail should stop.
The running machines were not touched: this only ever affected *new* installs.
- **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their
off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out
of one internal resource and, while looking perfectly healthy from outside, was accepting no
connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups
were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was
skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is
weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a
year** on the corrected measurement, not a cure.
- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The
second reinstall used to hit our own leftover; it was watched failing on the cycle that actually
fails, then watched passing.
- **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types
the code for an older set of backups, the machine now checks the packages we kept, recognises it, and
says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are
fine, write to us*. It deliberately promises no restore, because there is no button yet.
- **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown
banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302).
- **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere,
on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow
„Helyreállítási kód" and cost a real code.
- **The hub can see whether a machine's guest still has working networking** (R-319, first reader built
against R-264). A machine quietly repairing its own network over and over is now visible instead of
being a green tick; a machine that does not report it is drawn as unknown, never as healthy.
- **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an
account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word
from the name we retired last week, and false besides: it restores nothing. Five places, all in the
hub; no machine touched.
- **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down,
then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the
morning backup reminder had the same blind spot and is fixed with it.
- **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are
now checked for a retired name, and the guard has been watched catching one, ignoring an
explanation of one, and going quiet again.
- **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted.
## Broken, or knowingly incomplete
- **FIXED and proven on the machine: the off-site restore gives the app's data back** (R-354,
controller 0.218.0). The same planted files, the same steps, both runs on `demo-hp`: on the old
build the restore said „0 fájl visszaállítva", reported success, and the folder was simply not
there. On the new one it says „**0 fájl és 1 adatkötet visszaállítva**" and all five files come back
**byte for byte**, Hungarian accented names included. The message now names what came back, because
a restore that mentions only its file count is how a silent loss reads as a success. *(register:
R-354, CLOSED)*
- **FIXED and proven on the machine: Paperless's database is in the backup, and a restore of it now
takes an undo copy first** (R-355, controller 0.218.0). The dump was going into a folder named after
an app that does not exist, so nothing collected it — and because the same wrong name was used when
looking for the live database, a restore took **no undo copy at all**. Now: the dump is in the app's
own backup and in the off-site copy for the first time; the restore said „**0 fájl és 3 adatkötet és
az adatbázis visszaállítva**"; and with the undo deliberately made impossible the restore **refused
and did not even stop the app**. We no longer guess which app a database belongs to — Docker already
tells us. **One app of 53 was affected**, established with a check we first proved could catch a
planted second case. *(register: R-355, CLOSED)*
- **STILL BROKEN, and it is now the one that matters most: 40 of our 53 apps still cannot use the
off-site restore at all** (R-356). It refuses before it starts, says a running app „nincs telepítve"
— is not installed — and tells the customer to reinstall it "to the same place", which those apps
give them no way to choose. **Those are exactly the apps whose entire data is the thing R-354 just
fixed**, so today's fix cannot reach them until this one is done. *(register: R-356)*
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the
agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the
standard tool — which we never run. *(register: R-359)*
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
for a box we actually write to once a week. That volume is what turned a slow internal leak into
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
itself is untouched.** The honest health check is the resource count climbing, not the absence of an
alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too
optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years.
A deadline, not a comfort.
- **The leak is OURS, not Proxmox's, and asking fewer questions would not have fixed it** (R-344, found
2026-08-20). We finally looked at who was on the other end of the stuck connections. Every single one
belongs to **our own agent** on the two demo machines — it opens a connection to the off-site box on
each 15-minute cycle and never closes it, and neither does the box. The two Proxmox pollers that make
99.5% of the traffic leak **nothing at all**. So the plan recorded under R-336 — turn the question
rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed
fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a
deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)*
- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the
idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it
back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one
machine and left the other alone, and in the same hour the untouched one leaked 4 more connections
while the fixed one leaked none — with both doing exactly the same four rounds of work. **The
off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the
built-up connections released themselves when the agents restarted; the off-site box was only ever
read from, never touched. *(register: R-344)*
- **PUBLISHED the same day, on your word** (R-347, closed). Agent **0.130.0** is released, and the hub
now hands it to any new machine. Both demo machines run the exact published copy. Nothing else on
that screen was changed — in particular the controller floor was left alone, and the "minimum agent"
setting too, because raising that would have **stopped** machines getting updates rather than
helping them.
- **Two things the release itself turned up.** (1) The machines were briefly running a *different*
build of the same version number — harmless here, but nothing in the system would ever have noticed,
because everything compares the version *name*. Now corrected, and filed so it cannot repeat
(R-349). (2) **I printed the hub password into my own session log** while confirming the change
(R-350). It is not in git and not in any saved file — but it is in the log on this machine.
**Changing it is your call**; I can do it without ever showing the new one. Ask and I will.
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
installed the newer backup software for the practice, having first read its release notes and found
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
behaves exactly as before, which is the result the release notes predicted. **Two dated checks are
booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it.
- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was
still showing this morning's failure for four minutes after the copy was safely on the off-site box;
`demo-felhom` updated in under a minute. **It corrected itself** and both machines now read
correctly, so this is a note to watch, **not something broken** — but during the repair it looked
briefly like a second fault, which is the reason it is written down.
- **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both
machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was
not** — its control channel is still bound to the ordinary home network, which is exactly what that
change existed to stop. The machine works; the page is wrong, and it misled this session by an hour.
Your call which one to correct.
- **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine
belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a
half months and has been silent since 15 July, when its host record was deleted. There is no key, no
off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the
visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in
existence. Whether it stays parked is your call and is deliberately left open.
- **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built).
Today the honest answer is "your code is right, write to us" — and we can.
- **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine;
only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight.
- **The storage page has its own separate reason for showing an empty list** (R-298), untouched.
- **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent
update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others
are now recorded as deliberately unread, which is honest rather than fixed.
- **Two thirds of the standing picture is still unproven, and now you can ask** (R-326).
`python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of
those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine
is how many claims the 9 August review *lowered*, which is a different question.
- **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its
status may only be raised after the capability map moves first, which is a separate judgement.
## Working on next
**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two
demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
still untriaged against everything since.