diff --git a/REPORT.md b/REPORT.md index 638be722..79b228d7 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,92 +1,161 @@ -# REPORT — localisation slice 5: the catalog read path, the copy gate, and the pilot (R-560) +# REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0 -2026-09-20. Three repos touched: **felhom-controller v0.257.0** (the read path), -**app-catalog-felhom.eu** (freeze + gate + pilot), **felhom.eu** (this file, the architecture -document, the register, STATUS, the audit). Hub, agent, ISO and website untouched. +2026-09-21 (evening). Repos touched: **felhom.eu** (floor, docs, register, evidence), +**felhom-controller v0.261.0**, **app-catalog-felhom.eu** (two drill pairs, both reverted). +Architecture read first and named: `documentation/architecture/09-update-architecture.md` §3, §3b, +§4, §6.1, §6.2, §6.4, §8. -## Where it ended +--- -The operator read the pilot and said go, so **Part C shipped the same day**: the other fifty apps in -three pushes. **1 031 of the catalog's 1 032 customer-facing strings are English**, and the one that -is not is a Hungarian defect deliberately left to fall back (R-593). **The fleet floor was then -raised to 0.257.0**, as the operator asked, and the second demo box took it by itself in under twelve -seconds and rendered English. +## 1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session -**Nothing is waiting on the operator from this work.** - -## What was proven, and how it was seen - -Endpoint level. No browser exists on DooPlex, so every page was fetched with `curl` from inside the -guest at the controller's container address, with the customer `Host` header and a real logged-in -session — the same handlers and templates a browser drives, with only the drawing skipped. - -| claim | evidence | +| item | state | |---|---| -| The English app pages show English catalog text | demo-hp on 0.257.0: twelve apps opened individually across the pushes | -| **The English Apps list has NO Hungarian app text left** | all 53 apps on one page: the only Hungarian is the „Naprakész" badge (R-589) and the language picker naming itself | -| The floor delivered the FEATURE, not a version string | demo-felhom self-updated 0.255.0 → 0.257.0 in under 12 s, then rendered the English tagline | -| **The Hungarian did not move** | the same seven pages in Hungarian, before and after the catalog push: byte-identical apart from the per-session CSRF token — equal byte counts and equal hashes once normalised | -| **An old controller ignores the block** | demo-felhom on **0.255.0**: with the block synced (confirmed positively — the sync named the three apps and the block is in both the cache and the stack copy), pages hash identically before and after, and **no** parse warning in a **93-line** log window that contains the sync's own lines | -| The gate can convict | 33 decoy cases in the catalog repo, every one seen to convict or to pass as intended | -| The read path is not hollow | nine red-proofs in the controller, each seen to fail and then revert | +| **Part 0** floor to 0.260.0 | **done** | +| **Part 1** the three cuts | **done, but NOT as specified.** All three landed in `verifying`, never in `starting` — see §2. The brief allowed this explicitly and asked that it be said. | +| **Part 2** the lock + `reason` on the wire | **done**, five red-proofs, proven live | +| **Part 3** the unattended night | **done for the success night and the no-retry proof. The unattended HOLD was NOT produced** — see §4. | +| **Part 4** docs and rows | **done** | +| the caller script's ≤150-line budget | **167 lines.** Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. | +| Scenario C run by the measuring agent | **run by the coordinator instead.** The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. | +| a second drill bump/revert pair | **used.** The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (`ae08a037fd68`) added one real edge and one deliberately failing edge for Part 3. | -## What is still Hungarian on an English app page +**Two instrumentation failures of my own, recorded because they cost evidence:** +1. The first unattended run's stdout was piped through `tail`, which buffers, and the run was later + killed — **the caller's own log for the Scenario F press was lost.** The outcome survived on the + box; the log did not. The second run wrote straight to a file. +2. A background security review flagged the deliberately-broken `vikunja → alpine:3.20` catalog edge + as a supply-chain change. **It was right to.** Accepted deliberately — no customer or demo box runs + vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class + control — but the window is now closed by the revert, and it is named here rather than left in a + tool notification. -Measured with an accent scan AND an ASCII-folded stem scan, each with a positive and a negative -control. Three things; the first is correct. +--- -1. The language picker's own „Magyar" button — a picker names each language in its own tongue. -2. The update badge „Naprakész" and its tooltip — **R-589**. Slice 1 listed it; slice 2 was to take - it and closed without it. -3. The data-folder card's consequence sentence — **R-590**. It makes a promise about the customer's - files, and the label above it is already English, so the line reads half and half. +## 2. Claims in the brief that turned out wrong -The Apps list also shows fifty Hungarian descriptions: those are the apps nobody has translated yet. +**§2.4's open question — does `backupMgr.IsRunning()` cover the update's `backing-up` phase?** +**YES.** `RunAppBackupNow` calls `acquireRunning` (`internal/backup/update_guard.go:333`), so that one +phase was already protected. The gap was `checking`, `safety-dump`, `pinning`, `pulling`, `starting` +and `verifying`. **The live lock probe landed in `safety-dump`**, i.e. squarely in the previously +unprotected window rather than in the one that was already covered. Everything else §2.4 asserted +held at source. -## Claims in the plan that turned out wrong +**§2.3 — the resume path was READ, not measured. It is now measured**, three times, and it does what +it said. -| the plan said | measured | +**The stopped-guest byte path in my own brief to the measuring agent was wrong.** +`/var/lib/lxc/9202/rootfs/var/lib/felhom/...` is an empty mountpoint while 9202 is stopped, because +`mp0` is a separate raw volume; `pct mount 9202` does attach it. I had corrected one trap and +introduced a second. The positive control caught it. + +**"Four qualifying apps exist" — held.** vikunja, uptime-kuma, wishlist, glance: single-container, no +database sidecar, none on either demo box, all four target tags verified to exist upstream first. + +--- + +## 3. Part 0 — the floor + +Raised to **0.260.0** with MinAgent **0.131.0** declared (above the vouched golden 0.258.0, so the +declaration carries it — §3 decision 7). Hub log: + +``` +[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0") +[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0) +[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0) +``` + +Blast radius, read from the hub before saving: **3 boxes below** — `drill-r50` (0.213.0, BLOCKED), +`peti-felhom` (0.115.0, DOWN), `tester-1` (0.245.0, DOWN). None was reachable, so none moved; they +take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor. + +--- + +## 4. Part 1 — the three cuts (R-610, CLOSED) + +Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first. + +| | A — vikunja | B — uptime-kuma | C — wishlist | +|---|---|---|---| +| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 | +| cut | `pct stop` | `pct stop` | **controller container only** | +| phase at decision | `starting` | `verifying` | `starting` | +| cut latency | 3 759 ms | 3 016 ms | **1 675 ms** | +| phase it died in | `verifying` | `verifying` | `verifying` | +| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s | +| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s | +| four observables | **agree** | **agree** | **agree** | +| seeded data | **read back intact** | **read back intact** | not re-read (gap) | + +**The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.** +vikunja's own log: `Ran all migrations successfully` / `Vikunja version v2.6.0` at **12:28:26.881 +UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema +migration had already been applied to the customer's database when the power went. Recovery resumed +**forward**, so old-binary-on-migrated-database never happened — **but this branch is one step from +it**, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it. + +**Instrument limit:** `starting` lasts well under a second on this box. Three attempts, two cut +mechanisms, all landed in `verifying`. No phase was faked. `RecoverUpdates` handles `starting` and +`verifying` in **one branch**, so all three exercise the arm under test. A cut inside `starting` +itself needs an in-process fault injector. + +**Scenario C also showed the other apps were undisturbed** by the controller restart — uptime-kuma, +vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating +restarted. + +--- + +## 5. Part 2 — v0.261.0, proven live with a control + +See `felhom-controller/REPORT.md` for the code. The live proof is the part worth repeating: + +| probe | result | |---|---| -| "835 strings and the table is the complete list of copy fields" | **1 032** copy strings; the table omitted `deploy_fields[].placeholder` (13 of them) and counted only the strings with an accent | -| "832 with a Hungarian letter" | **832** ✔ | -| "three are ASCII-only Hungarian words — „Igen", „Nem", „Nincs"" | **~120** are ASCII-only Hungarian, and those three do not occur in this catalog at all. This is the one that chose an instrument: an accent-only gate passes „Aldomain" (53×) and „A szerver domain neve" (53×) inside an English block | -| "catalog copy reaches 13 templates" | **four**: `dashboard.html` (the app rows' second line), `stacks.html` (the Apps list), `app_info.html` (the app page, including the data-folder cards) and `deploy.html`. A fifth, `app_export.html`, reads `.Stack.Meta.Subdomain` — configuration, not copy. Five handler sites feed them | -| "no gate reads copy at all" | ✔ correct — none did | -| "the sync is 15 min with no pin" | ✔ correct; a manual trigger exists (`POST /api/sync`, the dashboard's "Sablonok frissítése") and was used | -| "`yaml.Unmarshal` is non-strict" | ✔ correct, and stronger than stated: this repo constructs **no** `yaml.Decoder` at all, so there is no `KnownFields` anywhere that could be turned on | -| the line numbers (`metadata.go` L15/L123/L158/L316) | ✔ all correct | -| "„Igen", „Nem", „Nincs" are the ASCII-only option labels" | the three ASCII-only option labels are **„Magyar", „Angol", „Magyar + Angol"** — language names, not yes/no words. „Igen" and „Nem" DO open vaultwarden's two option labels, but both of those carry accents later in the string and were already counted | +| manual self-update, **no** app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the **agent** refusal | +| manual self-update, app update **in flight** (`safety-dump`) | „Egy alkalmazás frissítése éppen folyamatban van…" — **our** refusal | -## Rows +**The sentence changed.** Guest 9202 has no host agent, so `TriggerUpdate` refuses either way — which +makes it the perfect negative control, because the new check sits *before* the agent check. Two +sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and +**restored to `false`** from a copy taken first; verified by re-reading the file. -**Opened:** R-589 (the update badge is Hungarian on an English page), R-590 (the data-folder card's -backup promise is Hungarian on an English page), R-591 (`Stack.Copy()` deep-copies five `Meta` fields -and not the new `I18n` map — safe today, which is why it is a row and not a fix), R-592 (three -defects inside the new catalog gate, **closed the same session**, each found by its own decoy rather -than by reading it), R-593 (papra describes a session-signing key as „the app's subdomain" — the one -string left untranslated), R-594 (the catalog gate can CONVICT a retrieval promise but has no way to -REGISTER a true one, which the shared vocabulary's own design calls for). +**The reverse direction — an app update refused while the controller swaps — is NOT staged live.** It +is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this +guest does not have. Stated as a gap. -**Closed:** R-560. +--- -Register: 281 → 287 rows. +## 6. Part 3 — the unattended night (R-611, CLOSED) -## Documents changed here +**Success:** `uptime-kuma` 2.5.0 → 2.5.1 applied with nobody pressing anything. -- `documentation/architecture/10-localisation.md` — §7 rewritten from **[DESIGN, proposed]** to - **[FACT] measured**, with the corrected numbers, the key-matching rules, and the gate's five - checks; new §10.6 for the slice. -- `documentation/backlog/OPEN-ITEMS.md` — the four rows above, and R-560's progress. -- `STATUS.md` — the operator note, in plain language, with the one read that is being asked for. -- `documentation/audits/i18n-slice5-2026-09-20/` — the captures from both boxes, the Hungarian parity - table, the English leftover scan with its controls, and a README saying what was NOT proven. +**No-retry:** after the revert left all four apps *ahead* of the catalog, the caller pressed each +**exactly once**, was refused `downgrade` (terminal), and pressed nothing across two further passes. +`never_again=['glance','uptime-kuma','vikunja','wishlist']`, `outcomes={}`. That is R-524 and R-609 +working together, unattended. -## Teardown, three layers +**The unattended HOLD was never produced, and the reason matters:** the only failing edge available +(`vikunja → alpine:3.20`) was **correctly refused by the within-a-major rule before it was ever +attempted**. The rule that makes automatic updates safe is the same rule that refuses the obvious way +to break one. Measuring it needs an image that passes the version test and still fails health. +**§3b Q4 therefore still rests on the ATTENDED hold from slice 4.** -- **Machine:** nothing installed or removed on either demo box; no app deployed; no VM created. Guest - 9201 on demo-hp was restarted onto the new controller image, which is the deploy itself. -- **Host:** the staged password file and the capture scripts were shredded from both Proxmox hosts and - both guests; `ls` confirms they are gone. -- **Hub:** nothing. No floor change, no golden, no vouch, no appliance registration. +--- -Both boxes left on Hungarian. +## 7. Rows, and the catalog + +**Closed:** R-608, R-609, R-610, R-611. **Opened:** R-612 (P1 — wishlist unusable on a fresh install +and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 — +stale update phase survives a redeploy). **Corrected:** R-520's closing pointer now names R-610. +Register 302 → **307**. + +**Catalog:** DRILL `573e41f5`, DRILL 2 `ae08a037`, **REVERT `f5f6a152`**. Every `image:` line in +`templates/` is byte-identical to the pre-drill `ff9717d3` — `git diff` over those paths is **0 +lines**. `catalog_since` reads 2026-09-21 on the four rather than the older dates, because the gate +requires an image move to carry the day's date in either direction and a revert is a move. + +**Teardown, three layers.** *Machine:* guest 9202 left running on v0.261.0 with the four throwaway +apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for +the remaining Q4 work and removing them would cost the next session the seeding. *Host:* demo-hp +untouched apart from 9202; **guest 9201 never touched.** *Hub:* nothing provisioned, nothing +enrolled; 9202 reports to no hub by design. `/tmp/.ctlpw` shredded. diff --git a/STATUS.md b/STATUS.md index 5e9bc859..ddb2cff1 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,63 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-21 (afternoon) — the update feature: a box can no longer be offered an older version as an "update", and I measured how far behind everything actually is.** +**Updated 2026-09-21 (evening) — I cut the power to a machine in the middle of an app update, three times, after the new version had already changed the data. It survived every time.** + +**The fleet version is now 0.260.0.** You approved it. Both demo machines have it. Three machines are +below it — the drill box, the tester's box and Peti's — and all three are switched off or blocked, so +they will take it when they come back. + +**The test the last session skipped without telling you: I ran it.** A machine updated an app with +**nobody pressing anything**, start to finish. It also proved the safer half: when the machine is told +to do something it must not do, it tries **once**, is refused, and then **never asks again**. Four +apps, three rounds, no repeats. + +**The dangerous power cut — the one nobody had measured.** The earlier test cut the power *before* the +new version started, which is the easy case. I cut it *after*. Three times, three apps, two different +ways of pulling the plug. **Every time the machine came back, finished the job, and told the truth.** +The version it thinks it runs, the version written down, and the version actually running all agreed, +every time. One app had already rewritten its own database when the power went — and the test data +read back unchanged. + +**One small lock added.** The machine used to update its own software at 04:30 without checking +whether it was busy updating an app. The window I proposed for automatic app updates covers 04:30, so +the two could have met. Now they queue politely behind each other. I proved it on a real machine: with +an app update running, the machine's own update is refused with a plain sentence; with nothing running, +it is not. **It never gets stuck** — an app that fails does not block the machine's own updates for ever. + +**Three faults found while doing this, all written down, none fixed today.** +1. **Wishlist cannot be signed up to on a fresh machine**, and the error it shows is *wrong* — it says + the account already exists when the real problem is that its first-time setup ran out of memory. + The machine reports the app healthy throughout. This is the one I would fix first. +2. **Uptime-Kuma sits on its setup screen with no login and no monitoring**, and the machine still + reports it healthy. A false green is the kind nothing ever catches. +3. A leftover "updated" label can survive an app being removed and reinstalled. + +**Two honest gaps.** I could never catch the cut in the very first instant of an update — it lasts +under a second, and three attempts with two methods all landed a moment later. And the unattended test +never produced a *stopped* app, because the safety rule correctly refused the broken test case before +it could be tried. Both are written down rather than glossed. + +**A security check flagged one of my test changes** — I briefly pointed an app in the catalog at a +dummy image to make it fail on purpose. It is the standard way to test a failure, no machine of yours +runs that app, and it is now reverted. Flagging it because you should hear it from me. + +**Rows.** Two closed, five opened, one pointer corrected. 307 in total. + +**Needs you — one thing, and it is smaller than last time.** + +**Raise the fleet version again, to 0.261.0?** That carries the lock to the rest of the machines. It is +on the scratch box only today. +**If you do nothing:** nothing breaks. The lock matters most once apps update themselves, which they +still do not. + +The seven questions from this morning are unchanged and still yours. Two of them now have measurements +beside them instead of guesses. + + +## Previous note + + +**Updated 2026-09-21 (afternoon, earlier) — the update feature: a box can no longer be offered an older version as an "update", and I measured how far behind everything actually is.** **A decision I took on my own — you can reverse it.** When we move an app back to an older version in our catalog, a machine that already took the newer one used to show **"Update available"** — and the diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index ea166136..300ea17c 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -227,6 +227,20 @@ window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy o new mechanism. **If nothing is decided:** Slice 6 cannot be built at all — every other question below is downstream of this one. +**⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF.** Found 2026-09-21 (R-608) by +reading the clock rather than by a failure: the controller self-updates daily at +`self_update.auto_update_time`, **default 04:30**, and again from `MaybeAutoUpdate` after ANY hub +report once a floor sits above the box — so **at any hour, not only at 04:30**. Either path restarts +the controller container, which is the supervisor of a running app update. + +**What v0.261.0 now guarantees, so this question can be answered without also solving that one:** the +two cannot overlap in either direction. The controller defers its own swap while a guarded app update +is in flight (retrying on the next report, exactly as it already did for a running backup), and +`UpdatePreflight` refuses `self_updating` while a swap is in progress. **The lock does not latch** — a +held app does not block the controller's updates for ever. So the window may contain 04:30; the two +jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the +operator prefers the box to take its own update first, that is a scheduling choice still open here. + ### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? *The button's rule and the automatic rule can differ. Should they?* @@ -289,6 +303,22 @@ major gets automated by accident. `reason: update_failed`), plus one mail. **If nothing is decided:** the safe default is no automatic update at all, because a hold nobody is told about is worse than a version nobody moved. +**MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured.** The +unattended night ran (`audits/update-arc-gaps-2026-09-21/09-unattended-night.md`). What it proved: +an app updates itself end to end with nobody pressing anything; one app takes **51 s – 1 m 26 s** +including the health wait; the caller needs **no new controller code**, only the existing guarded +Update plus `UpdateRefusal.Reason` on the wire (v0.261.0, R-609); and **a terminally-refused app is +pressed exactly ONCE and never again** — four apps, three passes, proven. + +**What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the +run.** The only failing edge available was `vikunja → alpine:3.20`, and the caller **correctly +refused to attempt it**: different repositories cannot be ordered, so the edge is "across" and +belongs to a human by decision 3. **The rule that makes automatic updates safe is the same rule that +refuses the obvious way to break one.** Measuring the unattended hold needs an edge that PASSES the +within-a-major test and still fails its health check — same repository, same major, a tag that +starts and does not serve — which probably means a purpose-built image rather than a catalog move. +**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).** + ### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? *Eleven templates, and the image performs no conversion — it refuses to start on an older major's @@ -682,7 +712,23 @@ window and the per-app switch, and calling `Manager.StartGuardedUpdate`. **It mu update's health wait is the defect that release fixed, and a second unattended caller is exactly the shape that finds it again. -**Ships behind `auto_update: off` with no UI until the operator answers Q1.** +**The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed** +(v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient** +— try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again +until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not +retried. **Before the reason reached the wire the only safe readings were "give up on everything" or +"press for ever"**, which is why this is listed as a dependency of the slice rather than a detail of +it. A working caller in this exact shape exists as evidence, not product: +`audits/update-arc-gaps-2026-09-21/unattended-caller.py`. + +**Ships behind `app_update.unattended: false` with no UI until the operator answers Q1.** + +⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.** +`self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default +**04:30** at L422) are the CONTROLLER's own update. An earlier draft of this section said Slice 6 +"ships behind `auto_update: off`"; two settings with that name, one meaning the controller and one +meaning apps, is the kind of collision that is only discovered by an operator who turned off the +wrong one. The app-scoped key is `app_update.*`. ### 6.3 Slice 7, as it would be built (OPEN — R-451; needs Q7) @@ -715,10 +761,10 @@ headlessly (R-460). | leg | what | cost | |---|---|---| | A | the **15 database services** — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus `adventurelog`'s postgis — one edge each, fixture per app | **15–25 CC-hours**, dominated by seed routes; ~30 min machine time at the median; ~25 GB | -| B | one **power cut mid-update** on a real version change, in `pulling` and again in `starting` | 1–2 CC-hours (R-520 — the first half is measured in this session) | +| B | ~~one power cut mid-update~~ **DONE 2026-09-21 (R-610)** — measured THREE times, two cut mechanisms, three apps: `pulling` (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. **What remains: a cut landing inside `starting` itself** (it lasts well under a second; needs an in-process fault injector, not a faster shell) | 0 — spent | | C | one **PostgreSQL `pg_upgrade` rehearsal**, the Q5 edge, on one app before any of the eleven | 3–4 CC-hours | | D | one **downgrade refusal** | **already done** — v0.260.0, proven live 2026-09-21 | -| E | the **automatic night** on a throwaway: one app, one real catalog step, inside a simulated window, with the guard and the hold; then the same edge made to fail → HOLD, the event, no retry loop | 2–3 CC-hours | +| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image | | F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised | **Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C diff --git a/documentation/audits/update-arc-gaps-2026-09-21/00-api-recipe.md b/documentation/audits/update-arc-gaps-2026-09-21/00-api-recipe.md new file mode 100644 index 00000000..db226c2e --- /dev/null +++ b/documentation/audits/update-arc-gaps-2026-09-21/00-api-recipe.md @@ -0,0 +1,112 @@ +# Driving the guest-9202 controller API from DooPlex — the working recipe (2026-09-21) + +Controller v0.260.0 in LXC guest **9202 `demo-hp-scratch`** on host `demo-hp`. + +## The one surprise that saves the most time + +**Guest 9202 is directly reachable from DooPlex on the home LAN at `192.168.0.114`.** +You do NOT have to `ssh demo-hp` + `pct exec` to drive the API — that was only needed for +`docker`/`pct`. Curling straight from DooPlex is what makes a 200 ms poll loop possible at all. + +Still mandatory: the **`Host: felhom.enkisfelhom.hu`** header (without it every route 404s with the +public page) and **`-k`** (self-signed cert on traefik). Port 443. `127.0.0.1:8080` inside the guest +is NOT open — traefik is the only door. + +## 0. The password, file→file, never through stdout + +```bash +python3 /mnt/5_hdd/felhom.eu/git/felhom.eu/scripts/read_credential.py PASSWORD /tmp/.ctlpw +chmod 600 /tmp/.ctlpw # expect 13 bytes; 15 means it is wearing its quotes +``` + +## 1. Log in and scrape the session CSRF (both are needed for any POST) + +curl's cookie jar drops `felhom_session`, so dump the headers and grep it out by hand. +The CSRF meta tag's closing quote must be stripped or you get a 65-char token that silently +mismatches — check the length is exactly 64. + +```bash +S=/tmp/ctl # any scratch dir +mkdir -p $S +B="https://192.168.0.114" +H='Host: felhom.enkisfelhom.hu' +PW=$(cat /tmp/.ctlpw) + +curl -sk -D $S/hdr.txt -o /dev/null -H "$H" -X POST --data-urlencode "password=$PW" "$B/login" +head -1 $S/hdr.txt # expect: HTTP/2 302 +grep -oiE 'felhom_session=[A-Za-z0-9._-]+' $S/hdr.txt | head -1 > $S/sess.txt # ~79 chars + +curl -sk -L -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/" -o $S/home.html +grep -oE ' $S/csrf.txt +echo "csrf len: $(wc -c < $S/csrf.txt)" # expect 65 = 64 + newline +``` + +## 2. The call helper — `c.sh GET /api/stacks` / `c.sh POST /api/sync '{...}'` + +```bash +cat > $S/c.sh <<'EOF' +#!/bin/bash +S=/tmp/ctl +B="https://192.168.0.114" +H='Host: felhom.enkisfelhom.hu' +M=$1; P=$2; D=$3 +SESS=$(cat $S/sess.txt); CT=$(cat $S/csrf.txt) +if [ "$M" = GET ]; then + curl -sk -H "$H" -H "Cookie: $SESS" "$B$P" +else + curl -sk -H "$H" -H "Cookie: $SESS" -H "X-CSRF-Token: $CT" \ + -H "Content-Type: application/json" -X "$M" ${D:+--data "$D"} "$B$P" +fi +EOF +chmod +x $S/c.sh +``` + +Worked endpoints (all verified today): + +| call | meaning | +|---|---| +| `c.sh GET /api/stacks` | every stack; `app_config.pinned_images`, `app_config.installed_images`, `template_images`, `catalog_images`, `updating`, `update_phase` | +| `c.sh GET /api/stacks/vikunja` | one stack, same shape — this is the 200 ms poll target | +| `c.sh POST /api/stacks//deploy '{"values":{"DOMAIN":"enkisfelhom.hu","SUBDOMAIN":"tasks"}}'` | real deploy path, answers 202 | +| `c.sh POST /api/stacks//update` | the Update button | +| `c.sh POST /api/sync` | pull the catalog git clone | +| `c.sh POST /api/stacks/rescan` | **always run this after a sync before reading any badge** (R-607) | +| `c.sh GET /api/stacks//logs?lines=200` | app container log | + +`pinned_images` and `installed_images` live under **`app_config`**, not at the top level — that +cost a few minutes. + +## 3. Reading the customer's app page, both languages + +```bash +curl -sk -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/app/vikunja" # Hungarian +curl -sk -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/app/vikunja?lang=en" # English +``` + +## 4. Shell into the guest (for docker / pct only) + +```bash +ssh -o StrictHostKeyChecking=accept-new demo-hp "pct exec 9202 -- bash -c ''" 2>/dev/null +``` +`2>/dev/null` drops the perl locale warnings. For anything with awkward quoting, pipe a script: + +```bash +cat <<'EOF' | ssh -o StrictHostKeyChecking=accept-new demo-hp \ + 'cat > /tmp/c.sh; pct push 9202 /tmp/c.sh /tmp/c.sh >/dev/null 2>&1; pct exec 9202 -- bash /tmp/c.sh' 2>/dev/null +docker ps --format '{{.Names}}\t{{.Image}}' +EOF +``` + +## 5. Paths inside the running guest + +- controller data dir (journal, catalog cache): + `/var/lib/docker/volumes/felhom-controller-data/_data/data/` + - `update-journal.json` — present only while an update is in flight + - `catalog-cache/` — a git clone; `git -C … log --oneline -1` tells you what the box actually has +- stacks: `/opt/docker/stacks//docker-compose.yml` (the LIVE rendered file) and `app.yaml` + +## 6. App front doors on 9202 (Host header per app, same IP) + +`tasks.` vikunja · `status.` uptime-kuma · `wishes.` wishlist · `dashboard.` glance — all +`…enkisfelhom.hu` against `https://192.168.0.114`. diff --git a/documentation/audits/update-arc-gaps-2026-09-21/01-baseline-deploy.txt b/documentation/audits/update-arc-gaps-2026-09-21/01-baseline-deploy.txt new file mode 100644 index 00000000..19f877f0 --- /dev/null +++ b/documentation/audits/update-arc-gaps-2026-09-21/01-baseline-deploy.txt @@ -0,0 +1,45 @@ +# 01 — baseline after deploying the four apps at today's pins +# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21T12:14:38Z + +== read at 2026-09-21T12:14:38.510303Z +glance state=running deployed=True deploying=False updating=False phase=- label=- + update_error=- + hold_reason =- + pinned ={"glance": "glanceapp/glance:v0.8.5"} + installed={"glance": {"ref": "glanceapp/glance:v0.8.5", "digest": "sha256:32ab73d80f2b8b5fb0735b0431deb36b93fbb6b2fb43592449b0178c8b83e350", "at": "2026-09-21T12:11:56Z"}} + template ={"glance": "glanceapp/glance:v0.8.5"} + catalog ={"glance": "glanceapp/glance:v0.8.5"} + health =healthy=True +uptime-kuma state=running deployed=True deploying=False updating=False phase=done label=Frissítve + update_error=- + hold_reason =- + pinned ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"} + installed={"uptime-kuma": {"ref": "louislam/uptime-kuma:2.4.0", "digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985", "at": "2026-09-21T12:11:45Z"}} + template ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"} + catalog ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"} + health =healthy=True +vikunja state=running deployed=True deploying=False updating=False phase=- label=- + update_error=- + hold_reason =- + pinned ={"vikunja": "vikunja/vikunja:2.3.0"} + installed={"vikunja": {"ref": "vikunja/vikunja:2.3.0", "digest": "sha256:f6b80393c1998cd5cd0dc38d24762c59ab4c10000a6f1032ef5b554e262cab93", "at": "2026-09-21T12:11:44Z"}} + template ={"vikunja": "vikunja/vikunja:2.3.0"} + catalog ={"vikunja": "vikunja/vikunja:2.3.0"} + health =healthy=True +wishlist state=running deployed=True deploying=False updating=False phase=- label=- + update_error=- + hold_reason =- + pinned ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"} + installed={"wishlist": {"ref": "ghcr.io/cmintey/wishlist:v0.66.0", "digest": "sha256:073ab4de0f27a93a79410172bedfa3947bed4718f050396547516def790f1f49", "at": "2026-09-21T12:12:17Z"}} + template ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"} + catalog ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"} + health =healthy=True + +## docker ps +wishlist ghcr.io/cmintey/wishlist:v0.66.0 Up 2 minutes (healthy) +glance glanceapp/glance:v0.8.5 Up 2 minutes (healthy) +uptime-kuma louislam/uptime-kuma:2.4.0 Up 2 minutes (healthy) +vikunja vikunja/vikunja:2.3.0 Up 2 minutes +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 Up About an hour (healthy) +filebrowser gtstef/filebrowser:1.3.3-stable Up About an hour (healthy) +traefik traefik:v3.6.7 Up About an hour diff --git a/documentation/audits/update-arc-gaps-2026-09-21/02-seeding.txt b/documentation/audits/update-arc-gaps-2026-09-21/02-seeding.txt new file mode 100644 index 00000000..5bc1f654 --- /dev/null +++ b/documentation/audits/update-arc-gaps-2026-09-21/02-seeding.txt @@ -0,0 +1,37 @@ +# 02 — seeding each app through its own front door, BEFORE any bump +# guest 9202, 2026-09-21T12:24:42Z + +## vikunja — seeded via its REST API (/api/v1/register, /login, PUT /projects, PUT /projects/2/tasks) +vikunja version: v2.3.0 +TASK 1 'SEED-VIKUNJA-CANARY-9f3c1e-20260921' desc= 'power-cut drill canary' +POSITIVE CONTROL: seed present = True +NEGATIVE CONTROL: absent string present = False + +## uptime-kuma — seeded via its own socket.io front door (the same API the browser uses) +# NOTE: uptime-kuma 2.4.0 first boot sits in the SETUP-DATABASE wizard ('Waiting for user action'). +# The controller reported the app running + healthy anyway. The wizard was passed through its own +# front door: POST /setup-database {"dbConfig":{"type":"sqlite"}} -> {"ok":true}. +MONITORS: ["SEED-KUMA-CANARY-4a91c7-20260921"] +POSITIVE CONTROL: seed present = true +NEGATIVE CONTROL: absent name present = false +read-back finished + +## wishlist — seeded via POST /signup then POST /lists//create-item +# NOTE: on first boot the image's own 'pnpm prisma db seed' was OOM-Killed at mem_limit 128M, +# so the Role/Group rows were missing and EVERY signup failed with a misleading +# 'User with username or email already exists' (real cause: FOREIGN KEY constraint). +# Repaired by running the image's own seed once (docker update --memory 512m, run seed, back to 128m). +wishlist version banner: 308 +list page code=200 +POSITIVE CONTROL occurrences of seed name: 2 +NEGATIVE CONTROL occurrences of absent name: 0 +context: '-[-->
SEED-WISHLIST-CANARY-7b2d44-20260921