0a5e9b14dc
gates / gates (push) Successful in 14s
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10.
152 lines
11 KiB
Markdown
152 lines
11 KiB
Markdown
# STATUS — what works, what's broken, what's next
|
||
|
||
**Updated 2026-08-18 (afternoon — a listening socket that served nobody, a rehearsal upgrade that changed
|
||
nothing, and a new install that is finally current).**
|
||
|
||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||
> If it does not fit, it belongs in the register instead.
|
||
|
||
## Waiting on you
|
||
|
||
*Nothing. All three questions that stood here were answered on 12–13 August and have moved to
|
||
**Decided** below. A decided question left in the deciding list is how a person loses track of what is
|
||
actually waiting.*
|
||
|
||
## Decided — and what would reopen each
|
||
|
||
*A decision with no trigger becomes a permanent silence, so each one names what would make us look
|
||
again.*
|
||
|
||
- **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a
|
||
support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks —
|
||
one request from a person who is not us. *(register: R-312)*
|
||
- **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the
|
||
only state in existence where a set-aside store is present and cannot be opened, which is the case
|
||
any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is
|
||
abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)*
|
||
- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is
|
||
unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed
|
||
happening outside a constructed test. *(register: R-303)*
|
||
|
||
## What works
|
||
|
||
Both demo machines are home, healthy and reporting on the approved pair — **controller 0.214.0, agent
|
||
0.129.0**, delivered by the floor rather than by hand. Off-site is credentialed on `demo-hp` and its
|
||
repository still opens with the machine's own key. `drill-r50` is reverted to `virgin`, powered off.
|
||
|
||
**What the fleet actually is, because two summaries have now been misread:** the hub holds **five
|
||
customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both
|
||
disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom`
|
||
is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a
|
||
record with no machine** — created 13 August, no host, no backups, nothing to lose.
|
||
|
||
## Shipped
|
||
|
||
- **The auto-update floor is current again** (R-343). You raised it to today's version this
|
||
afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*,
|
||
after the image is vouched, because it acts within seconds. It did: **one machine updated itself
|
||
nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not
|
||
paperwork.
|
||
- **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the
|
||
19th", that date is now read by the build system, and a push is refused once it passes. **It is not
|
||
a reminder service** — it speaks on the next push, not on the day — and that limit is written into
|
||
the check itself.
|
||
|
||
- **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image
|
||
had been two releases behind since the 14th — anyone installing would have received a version
|
||
missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked
|
||
and published, and **you vouched it**, which was the half that could not be done without you. **The
|
||
build system is green again for the first time since 14 August**, so the failure mail should stop.
|
||
The running machines were not touched: this only ever affected *new* installs.
|
||
|
||
- **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their
|
||
off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out
|
||
of one internal resource and, while looking perfectly healthy from outside, was accepting no
|
||
connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups
|
||
were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was
|
||
skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is
|
||
weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question
|
||
about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a
|
||
year** on the corrected measurement, not a cure.
|
||
|
||
- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The
|
||
second reinstall used to hit our own leftover; it was watched failing on the cycle that actually
|
||
fails, then watched passing.
|
||
- **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types
|
||
the code for an older set of backups, the machine now checks the packages we kept, recognises it, and
|
||
says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are
|
||
fine, write to us*. It deliberately promises no restore, because there is no button yet.
|
||
- **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown
|
||
banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302).
|
||
- **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere,
|
||
on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow
|
||
„Helyreállítási kód" and cost a real code.
|
||
- **The hub can see whether a machine's guest still has working networking** (R-319, first reader built
|
||
against R-264). A machine quietly repairing its own network over and over is now visible instead of
|
||
being a green tick; a machine that does not report it is drawn as unknown, never as healthy.
|
||
- **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an
|
||
account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word
|
||
from the name we retired last week, and false besides: it restores nothing. Five places, all in the
|
||
hub; no machine touched.
|
||
- **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down,
|
||
then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the
|
||
morning backup reminder had the same blind spot and is fixed with it.
|
||
- **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are
|
||
now checked for a retired name, and the guard has been watched catching one, ignoring an
|
||
explanation of one, and going quiet again.
|
||
- **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted.
|
||
|
||
## Broken, or knowingly incomplete
|
||
|
||
- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day,
|
||
for a box we actually write to once a week. That volume is what turned a slow internal leak into
|
||
last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak
|
||
itself is untouched.** The honest health check is the resource count climbing, not the absence of an
|
||
alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too
|
||
optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years.
|
||
A deadline, not a comfort.
|
||
- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we
|
||
installed the newer backup software for the practice, having first read its release notes and found
|
||
**nothing** about the fault we have. The update went cleanly and everything works, but the leak
|
||
behaves exactly as before, which is the result the release notes predicted. **Two dated checks are
|
||
booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it.
|
||
- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was
|
||
still showing this morning's failure for four minutes after the copy was safely on the off-site box;
|
||
`demo-felhom` updated in under a minute. **It corrected itself** and both machines now read
|
||
correctly, so this is a note to watch, **not something broken** — but during the repair it looked
|
||
briefly like a second fault, which is the reason it is written down.
|
||
- **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both
|
||
machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was
|
||
not** — its control channel is still bound to the ordinary home network, which is exactly what that
|
||
change existed to stop. The machine works; the page is wrong, and it misled this session by an hour.
|
||
Your call which one to correct.
|
||
|
||
- **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine
|
||
belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a
|
||
half months and has been silent since 15 July, when its host record was deleted. There is no key, no
|
||
off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the
|
||
visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in
|
||
existence. Whether it stays parked is your call and is deliberately left open.
|
||
- **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built).
|
||
Today the honest answer is "your code is right, write to us" — and we can.
|
||
- **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine;
|
||
only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight.
|
||
- **The storage page has its own separate reason for showing an empty list** (R-298), untouched.
|
||
- **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent
|
||
update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others
|
||
are now recorded as deliberately unread, which is honest rather than fixed.
|
||
- **Two thirds of the standing picture is still unproven, and now you can ask** (R-326).
|
||
`python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of
|
||
those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine
|
||
is how many claims the 9 August review *lowered*, which is a different question.
|
||
- **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its
|
||
status may only be raised after the capability map moves first, which is a separate judgement.
|
||
|
||
## Working on next
|
||
|
||
`demo-hp` is yours this evening — **this session did not touch it**. After that: the three remaining
|
||
R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent);
|
||
R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292),
|
||
still untriaged against everything since.
|