877fcd2a38
gates / gates (push) Successful in 16s
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with a negative control first — the same planted, hash-recorded fixture run through the same steps on both builds. R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the off-site copy or the restore; and because the same wrong name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label. Sweep proven able to convict before its count was trusted: one affected app of 53. R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit, before the database and inside the stopped window, and VolumesReplayed reaches the sentence. The half-false comment beside the skip is corrected and the half that still holds is named. Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b, verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are the operator's decision, and raising the floor is what puts this on demo-felhom, which is still on 0.217.0 and still has both defects. R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them (an existing guard), they are adoptable by hand, and doing it automatically would be a migration. Ceiling R-366 -> R-367.
219 lines
17 KiB
Markdown
219 lines
17 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
||
|
||
**Updated 2026-08-22 — both of last night's worst findings are fixed and proven on the machine.**
|
||
|
||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||
> If it does not fit, it belongs in the register instead.
|
||
|
||
## Waiting on you
|
||
|
||
*Nothing. All three questions that stood here were answered on 12–13 August and have moved to
|
||
**Decided** below. A decided question left in the deciding list is how a person loses track of what is
|
||
actually waiting.*
|
||
|
||
## Decided — and what would reopen each
|
||
|
||
*A decision with no trigger becomes a permanent silence, so each one names what would make us look
|
||
again.*
|
||
|
||
- **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a
|
||
support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks —
|
||
one request from a person who is not us. *(register: R-312)*
|
||
- **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the
|
||
only state in existence where a set-aside store is present and cannot be opened, which is the case
|
||
any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is
|
||
abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)*
|
||
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is
|
||
unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed
|
||
happening outside a constructed test. *(register: R-303)*
|
||
|
||
## What works
|
||
|
||
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.217.0, agent
|
||
0.130.0**. On `demo-hp` the floor delivered it unaided on 21 August: 0.216.0 → 0.217.0, **17 seconds**
|
||
from the operator pressing save to the new controller reporting healthy, with no customer action.
|
||
Off-site is credentialed on `demo-hp`, its repository opens with the machine's own key, and
|
||
`restic check` over the whole store reports **no errors**. `drill-r50` is reverted to `virgin`, powered
|
||
off. **`demo-hp` was reinstalled on 21 August and its off-site backup had been silent since 9 August**
|
||
— the rebuild lost the off-site target, and after that was healed every per-app off-site switch was
|
||
still off, so the nightly run reported "backup OK" having backed up nothing. Both are on again.
|
||
|
||
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
|
||
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
|
||
disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom`
|
||
is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a
|
||
record with no machine** — created 13 August, no host, no backups, nothing to lose.
|
||
|
||
## Shipped
|
||
|
||
- **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a
|
||
completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks
|
||
only ever watched how full a store was getting, and a failed reading was written to a log nobody
|
||
reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the
|
||
weekly backup that happened to fall inside it. After about half an hour of being unable to see a
|
||
store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back.
|
||
**Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have
|
||
caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That
|
||
second check is written down as the next step.
|
||
|
||
- **The auto-update floor is current again** (R-343). You raised it to today's version this
|
||
afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*,
|
||
after the image is vouched, because it acts within seconds. It did: **one machine updated itself
|
||
nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not
|
||
paperwork.
|
||
- **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the
|
||
19th", that date is now read by the build system, and a push is refused once it passes. **It is not
|
||
a reminder service** — it speaks on the next push, not on the day — and that limit is written into
|
||
the check itself.
|
||
|
||
- **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image
|
||
had been two releases behind since the 14th — anyone installing would have received a version
|
||
missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked
|
||
and published, and **you vouched it**, which was the half that could not be done without you. **The
|
||
build system is green again for the first time since 14 August**, so the failure mail should stop.
|
||
The running machines were not touched: this only ever affected *new* installs.
|
||
|
||
- **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their
|
||
off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out
|
||
of one internal resource and, while looking perfectly healthy from outside, was accepting no
|
||
connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups
|
||
were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was
|
||
skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is
|
||
weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question
|
||
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a
|
||
year** on the corrected measurement, not a cure.
|
||
|
||
- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The
|
||
second reinstall used to hit our own leftover; it was watched failing on the cycle that actually
|
||
fails, then watched passing.
|
||
- **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types
|
||
the code for an older set of backups, the machine now checks the packages we kept, recognises it, and
|
||
says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are
|
||
fine, write to us*. It deliberately promises no restore, because there is no button yet.
|
||
- **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown
|
||
banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302).
|
||
- **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere,
|
||
on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow
|
||
„Helyreállítási kód" and cost a real code.
|
||
- **The hub can see whether a machine's guest still has working networking** (R-319, first reader built
|
||
against R-264). A machine quietly repairing its own network over and over is now visible instead of
|
||
being a green tick; a machine that does not report it is drawn as unknown, never as healthy.
|
||
- **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an
|
||
account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word
|
||
from the name we retired last week, and false besides: it restores nothing. Five places, all in the
|
||
hub; no machine touched.
|
||
- **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down,
|
||
then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the
|
||
morning backup reminder had the same blind spot and is fixed with it.
|
||
- **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are
|
||
now checked for a retired name, and the guard has been watched catching one, ignoring an
|
||
explanation of one, and going quiet again.
|
||
- **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted.
|
||
|
||
## Broken, or knowingly incomplete
|
||
|
||
- **FIXED and proven on the machine: the off-site restore gives the app's data back** (R-354,
|
||
controller 0.218.0). The same planted files, the same steps, both runs on `demo-hp`: on the old
|
||
build the restore said „0 fájl visszaállítva", reported success, and the folder was simply not
|
||
there. On the new one it says „**0 fájl és 1 adatkötet visszaállítva**" and all five files come back
|
||
**byte for byte**, Hungarian accented names included. The message now names what came back, because
|
||
a restore that mentions only its file count is how a silent loss reads as a success. *(register:
|
||
R-354, CLOSED)*
|
||
- **FIXED and proven on the machine: Paperless's database is in the backup, and a restore of it now
|
||
takes an undo copy first** (R-355, controller 0.218.0). The dump was going into a folder named after
|
||
an app that does not exist, so nothing collected it — and because the same wrong name was used when
|
||
looking for the live database, a restore took **no undo copy at all**. Now: the dump is in the app's
|
||
own backup and in the off-site copy for the first time; the restore said „**0 fájl és 3 adatkötet és
|
||
az adatbázis visszaállítva**"; and with the undo deliberately made impossible the restore **refused
|
||
and did not even stop the app**. We no longer guess which app a database belongs to — Docker already
|
||
tells us. **One app of 53 was affected**, established with a check we first proved could catch a
|
||
planted second case. *(register: R-355, CLOSED)*
|
||
- **STILL BROKEN, and it is now the one that matters most: 40 of our 53 apps still cannot use the
|
||
off-site restore at all** (R-356). It refuses before it starts, says a running app „nincs telepítve"
|
||
— is not installed — and tells the customer to reinstall it "to the same place", which those apps
|
||
give them no way to choose. **Those are exactly the apps whose entire data is the thing R-354 just
|
||
fixed**, so today's fix cannot reach them until this one is done. *(register: R-356)*
|
||
- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the
|
||
agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the
|
||
standard tool — which we never run. *(register: R-359)*
|
||
|
||
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
|
||
for a box we actually write to once a week. That volume is what turned a slow internal leak into
|
||
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
|
||
itself is untouched.** The honest health check is the resource count climbing, not the absence of an
|
||
alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too
|
||
optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years.
|
||
A deadline, not a comfort.
|
||
- **The leak is OURS, not Proxmox's, and asking fewer questions would not have fixed it** (R-344, found
|
||
2026-08-20). We finally looked at who was on the other end of the stuck connections. Every single one
|
||
belongs to **our own agent** on the two demo machines — it opens a connection to the off-site box on
|
||
each 15-minute cycle and never closes it, and neither does the box. The two Proxmox pollers that make
|
||
99.5% of the traffic leak **nothing at all**. So the plan recorded under R-336 — turn the question
|
||
rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed
|
||
fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a
|
||
deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)*
|
||
- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the
|
||
idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it
|
||
back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one
|
||
machine and left the other alone, and in the same hour the untouched one leaked 4 more connections
|
||
while the fixed one leaked none — with both doing exactly the same four rounds of work. **The
|
||
off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the
|
||
built-up connections released themselves when the agents restarted; the off-site box was only ever
|
||
read from, never touched. *(register: R-344)*
|
||
- **PUBLISHED the same day, on your word** (R-347, closed). Agent **0.130.0** is released, and the hub
|
||
now hands it to any new machine. Both demo machines run the exact published copy. Nothing else on
|
||
that screen was changed — in particular the controller floor was left alone, and the "minimum agent"
|
||
setting too, because raising that would have **stopped** machines getting updates rather than
|
||
helping them.
|
||
- **Two things the release itself turned up.** (1) The machines were briefly running a *different*
|
||
build of the same version number — harmless here, but nothing in the system would ever have noticed,
|
||
because everything compares the version *name*. Now corrected, and filed so it cannot repeat
|
||
(R-349). (2) **I printed the hub password into my own session log** while confirming the change
|
||
(R-350). It is not in git and not in any saved file — but it is in the log on this machine.
|
||
**Changing it is your call**; I can do it without ever showing the new one. Ask and I will.
|
||
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
|
||
installed the newer backup software for the practice, having first read its release notes and found
|
||
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
|
||
behaves exactly as before, which is the result the release notes predicted. **Two dated checks are
|
||
booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it.
|
||
- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was
|
||
still showing this morning's failure for four minutes after the copy was safely on the off-site box;
|
||
`demo-felhom` updated in under a minute. **It corrected itself** and both machines now read
|
||
correctly, so this is a note to watch, **not something broken** — but during the repair it looked
|
||
briefly like a second fault, which is the reason it is written down.
|
||
- **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both
|
||
machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was
|
||
not** — its control channel is still bound to the ordinary home network, which is exactly what that
|
||
change existed to stop. The machine works; the page is wrong, and it misled this session by an hour.
|
||
Your call which one to correct.
|
||
|
||
- **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine
|
||
belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a
|
||
half months and has been silent since 15 July, when its host record was deleted. There is no key, no
|
||
off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the
|
||
visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in
|
||
existence. Whether it stays parked is your call and is deliberately left open.
|
||
- **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built).
|
||
Today the honest answer is "your code is right, write to us" — and we can.
|
||
- **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine;
|
||
only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight.
|
||
- **The storage page has its own separate reason for showing an empty list** (R-298), untouched.
|
||
- **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent
|
||
update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others
|
||
are now recorded as deliberately unread, which is honest rather than fixed.
|
||
- **Two thirds of the standing picture is still unproven, and now you can ask** (R-326).
|
||
`python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of
|
||
those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine
|
||
is how many claims the 9 August review *lowered*, which is a different question.
|
||
- **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its
|
||
status may only be raised after the capability map moves first, which is a separate judgement.
|
||
|
||
## Working on next
|
||
|
||
**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two
|
||
demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining
|
||
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
|
||
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
|
||
still untriaged against everything since.
|