R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step.
This commit is contained in:
@@ -1,80 +1,77 @@
|
|||||||
# REPORT — DRILL R-356b: the off-site restore for a driveless app that HAS a database (2026-08-22)
|
# REPORT — R-379/R-380/R-381/R-382: the undo copy goes back (2026-08-22)
|
||||||
|
|
||||||
**A drill, not an implementation.** No production code was written, no version bumped, no CHANGELOG
|
Companion to `felhom-controller` **v0.220.0 → v0.220.1 → v0.220.2**. Full record:
|
||||||
entry made. The deliverables are a findings document, four register rows and a capability-map update.
|
`documentation/audits/DRILL-r379-rollback-2026-08-22/`.
|
||||||
|
|
||||||
Full record: `documentation/audits/DRILL-r356b-driveless-db-restore-2026-08-22/`
|
## What shipped
|
||||||
|
|
||||||
## What was measured
|
**R-379 and R-380 were one failure with one fix.** Both ended with a half-restored database; the only
|
||||||
|
difference was whether it looked broken. When the replay fails, the product now re-applies the
|
||||||
|
customer's own pre-restore copy — the same `ImportDump` call a person ran by hand yesterday to recover
|
||||||
|
both apps — and the app comes back with a message saying **both** that the restore failed and that the
|
||||||
|
data is as it was.
|
||||||
|
|
||||||
Ten of the forty driveless apps carry a database. **I re-measured that count myself and got 10** — the
|
**The whole undo set, matched on the run's own stamp**, never on the `pre-restore-` prefix and never
|
||||||
same ten the runbook names. For those ten, restoring is a five-leg operation that, until this week,
|
just the first file. **When the rollback also fails the app is held stopped** — the operator's ruling —
|
||||||
never ran at all: R-356 refused before any of it started.
|
with every start path refusing it, the app-stop marker ended so nothing auto-restarts it, and the row
|
||||||
|
red rather than green.
|
||||||
|
|
||||||
Both engines were walked end to end on `demo-hp`: `docmost` (Postgres 16) and `bookstack`
|
**R-381:** the failure message stopped pasting engine output (407→257 bytes on Postgres; the MariaDB
|
||||||
(MariaDB 12.3), each deployed for this drill, planted through the app's **own** interface, destroyed
|
one had been 615 bytes with rows out of the customer's own database). The full text now reaches the
|
||||||
for real, and restored through the exact endpoint the UI's button posts to.
|
operator log, which never had it.
|
||||||
|
**R-382:** the summary log prints the volume count it already held.
|
||||||
|
**Also:** undo copies resolve to their own app and are capped at 3.
|
||||||
|
|
||||||
## The three answers
|
## Documents updated here
|
||||||
|
|
||||||
**Q1 — does it complete? YES.** All five legs ran in order and all succeeded — 32 s for Postgres,
|
- `documentation/architecture/07-backup-architecture.md` §6.3 — a dated **[DESIGN]** paragraph on the
|
||||||
25 s for MariaDB. Data back, apps healthy, accented names byte-identical in both directions.
|
failure ladder **replay → rollback → hold**, and why an engine flag does not close it.
|
||||||
|
- `STATUS.md` — the outcome in plain words; the deciding section says what happens if nothing is done.
|
||||||
|
- `documentation/backlog/` — R-379…R-382 compressed into `CLOSED-ITEMS.md`, each keeping its title,
|
||||||
|
shipping version, evidence path and every sentence that states a rule. Full text:
|
||||||
|
`git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md`.
|
||||||
|
|
||||||
**Q2 — which leg returned the data? The SQL dump.** A three-way discriminator (volume tar
|
**Register size:** `OPEN-ITEMS.md` **330 683 → 325 236 bytes**; `CLOSED-ITEMS.md` **63 507 → 66 777**.
|
||||||
`ORIGINAL-VALUE-A`, altered dump `ALTERED-VALUE-B`, live `LIVE-VALUE-C3`) returned **`ALTERED-VALUE-B`**.
|
|
||||||
The ordering the code comment asserts holds in practice. This **confirms R-164's F17 claim on a second
|
|
||||||
path** — R-164 cites `restore_unit.go`, the local restore; this measures `offbox_reconstitute.go`.
|
|
||||||
The mutation was applied to the prepared scratch only, and the store was proved unmutated afterwards
|
|
||||||
by re-preparing a fresh scratch (sha256 back to `c5414f24…`).
|
|
||||||
|
|
||||||
**Q3 — does a failure tell the truth? Partly.** The customer does see a failure and the undo copy is
|
## The live walk found two defects in the fix itself
|
||||||
named. But two things are wrong, and they are the drill's findings.
|
|
||||||
|
|
||||||
## Findings filed — R-379 … R-382
|
**Both are recorded because the walk, not the tests, caught them.**
|
||||||
|
|
||||||
- **R-379 (HIGH)** — the undo copy is valid, is named, and **nothing in the product can apply it**.
|
1. **The rollback used a dead container (fixed v0.220.1).** The DB-only start re-creates the DB
|
||||||
Proven by applying it by hand on both engines and getting the exact prior state back.
|
container, so the id captured at dump time is dead by rollback time. Measured: captured
|
||||||
`pre-restore-` files are deliberately skipped at three code sites; the filename appears only inside
|
`9adbc14f9af6`, re-created `309795897b82`, rollback timed out after 30 s — **the app was held for an
|
||||||
an error string.
|
infrastructure reason while its data was recoverable.** No unit test could see it: they all inject
|
||||||
- **R-380 (HIGH)** — a failed **MariaDB** replay leaves a partially-applied database behind an app
|
the import seam and never look at container identity.
|
||||||
reporting `health=healthy, running=true, restarts=0`. `bookstack`'s schema-version ledger was wiped
|
2. **The operator route did not take effect (fixed v0.220.2).** `--clear-restore-hold` runs as a second
|
||||||
to 0 rows while its user data stayed intact and the dashboard said fine. Postgres, by contrast,
|
process; it cleared the file and the running controller went on refusing. Found by using it.
|
||||||
fails visibly (crash-loop). H3 fired — but not in its predicted shape: the prediction was a *quiet
|
|
||||||
success*; what happens is a loud error and a silent inconsistency.
|
|
||||||
- **R-381 (MEDIUM)** — the failure message pastes raw engine stderr into the Hungarian customer
|
|
||||||
surface: 407 bytes for Postgres, **615 for MariaDB, whose middle is an `INSERT INTO migrations
|
|
||||||
VALUES (…)` listing — actual table rows shown to the customer.**
|
|
||||||
- **R-382 (LOW)** — the reconstitution's summary log omits the volume count it already has. The
|
|
||||||
customer-facing flash names the volumes; the operator log does not.
|
|
||||||
|
|
||||||
**Register: 325 236 bytes before, 330 683 after.** Ceiling was R-378; next free id is now R-383.
|
## A red-proof that PASSED
|
||||||
|
|
||||||
## Also recorded
|
Of nine mutations, **one did not fail its test** and is reported rather than omitted: the R-381
|
||||||
|
behavioural test injected below `ImportDump`, so a leak reintroduced inside `ImportDump` was invisible
|
||||||
|
to it. A guard now sits at that layer and the mutation convicts.
|
||||||
|
|
||||||
- **R-361 reproduced independently** on a second app: after the first reconstitution docmost's
|
## What did not reproduce
|
||||||
`db-dumps/` held only `pre-restore-*` files. Not re-filed — noted as corroboration.
|
|
||||||
- **`restic check` passed** at the end: `no errors were found`, 29 snapshots.
|
|
||||||
- **The `-db` suffix attribution is correct** for `bookstack-db` — the R-355 shape does not reproduce.
|
|
||||||
- **Observed, not filed:** a newly deployed app is absent from the off-site set until switched on by
|
|
||||||
hand. Plausibly deliberate; the consequence is stated so the default can be judged.
|
|
||||||
- **A flaw in the drill's own method, recorded rather than hidden:** the first accented title was
|
|
||||||
double-escaped by a shell chain and stored as literal ASCII. Caught by reading the stored bytes back
|
|
||||||
as hex, and re-measured properly in Phase 1b.
|
|
||||||
|
|
||||||
## Capability map
|
The undo copies rendering as app rows on the customer's backup page. The live page was read **before**
|
||||||
|
any change: zero `pre-restore` strings while four such files sat on disk, with a positive control
|
||||||
|
showing 8 real rows. The phantom name was real as a map **key**, never a row. Fixed as a naming defect;
|
||||||
|
their visibility is unchanged and deliberate.
|
||||||
|
|
||||||
The 2026-08-21 narrowing of *"A customer's file survives a machine rebuild and comes back"* is now
|
## The golden was baked in this session, and why
|
||||||
history: both defects it named (R-354, R-356) are closed and proven. The row records what is now
|
|
||||||
walked — including this drill — and states plainly what is still **not** claimed: the success path is
|
|
||||||
proven, the recovery-from-a-bad-restore path is not.
|
|
||||||
|
|
||||||
## Teardown, three layers
|
The `golden-currency` gate refused this docs push: 0.220.2 was released with no golden. **That block
|
||||||
|
is not circular** — a golden needs the controller image, which was already pushed, not this commit —
|
||||||
|
so the gate was satisfied by doing the work it asked for rather than bypassed with `--no-verify`.
|
||||||
|
**No push in this session used `--no-verify`.**
|
||||||
|
|
||||||
1. `docmost` and `bookstack` were deployed by this drill and are **RETAINED** with their planted data —
|
Golden **0.220.2**, sha256 `cb439418c7005ce01bcb8126bb6385688c2f408c3c4649f6740001bc2c864ed5`,
|
||||||
it is the evidence, and they are the only deployed members of this app class on the box. Both left
|
657 271 965 B, round-trip verified. All five markers hit, both negative controls at zero, both
|
||||||
healthy and sane.
|
token-leak greps proved able to convict before their zeros were accepted. Record:
|
||||||
2. No `pvesm` "before" snapshot was taken — **said plainly rather than reconstructed.** Measured
|
`documentation/tests/golden-0.220.2-2026-08-22/`.
|
||||||
directly: ~233 MB of volumes inside guest 9201 (16 % of 69 GB used).
|
|
||||||
3. **No hub-side record was created.** No customer, no appliance. Nothing to dispose of.
|
|
||||||
|
|
||||||
All phases were run. Nothing was dropped.
|
## Operator follow-up
|
||||||
|
|
||||||
|
**Vouch** the golden — Hub → Configuration → Day-0 artifacts, a **three-field** save:
|
||||||
|
`golden_version` **0.220.2**, `agent_version` **0.130.0**, `min_agent` **0.129.0**. **Then** raise the
|
||||||
|
floor to **0.220.2**, last, in its own save.
|
||||||
|
|||||||
@@ -12,16 +12,14 @@ NOT yet delivered: two steps below are yours.**
|
|||||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||||
nothing.*
|
nothing.*
|
||||||
|
|
||||||
1. **Vouch the golden carrying controller 0.219.0** — Hub → Configuration → Day-0 artifacts.
|
1. **Vouch the golden carrying controller 0.220.2** — Hub → Configuration → Day-0 artifacts.
|
||||||
**It is baked, published and round-trip verified** (`documentation/tests/golden-0.219.0-2026-08-22/`);
|
**It is already baked, published and round-trip verified**
|
||||||
only the vouch is left, and only you can do it. **It is a THREE-field save, not one:**
|
(`documentation/tests/golden-0.220.2-2026-08-22/`); only the vouch is left, and only you can do it.
|
||||||
`golden_version` → **0.219.0**, `agent_version` → **0.130.0**, `min_agent` → **0.129.0**. Moving
|
**It is a THREE-field save:** `golden_version` → **0.220.2**, `agent_version` → **0.130.0**,
|
||||||
`golden_version` alone ships this controller onto an agent older than it declares it needs.
|
`min_agent` → **0.129.0**. **Then** raise the floor to **0.220.2**, last, in its own save.
|
||||||
**If you do nothing:** a machine installed today still receives 0.218.0 — the image exists, on the
|
**If you do nothing:** the fleet stays on 0.219.0, so a failed database restore still leaves an app
|
||||||
shelf, undelivered. Reversible: re-select the old values and Save.
|
broken with an unusable copy — the thing today's release fixes reaches nobody. New machines still
|
||||||
2. **Then raise the auto-update floor to 0.219.0 — last, in a separate save.** It acts within seconds.
|
receive 0.219.0. The build system stays red about it and will mail you on every push.
|
||||||
**If you do nothing:** every existing machine stays on 0.218.0, so the fix below reaches nobody and
|
|
||||||
40 of 53 apps stay un-restorable on the actual fleet. *(register: R-343's rule)*
|
|
||||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||||
@@ -52,6 +50,19 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
|||||||
|
|
||||||
## Shipped
|
## Shipped
|
||||||
|
|
||||||
|
- **A failed database restore now puts your data back by itself** (R-379/R-380, controller 0.220.2,
|
||||||
|
proven on `demo-hp`). Until today, if a restore of an app's database went wrong, the machine had
|
||||||
|
already taken a copy of your live database — a good copy — and **nothing in the product could put it
|
||||||
|
back.** You were shown a filename. On one of the two database types it was worse: part of the
|
||||||
|
restore applied, part did not, and **the dashboard said the app was healthy**. Now the machine puts
|
||||||
|
your own copy back automatically and says plainly: the restore failed, your data is as it was, the
|
||||||
|
app is running. Proven on both database types, byte-identical both times.
|
||||||
|
**If even that fails**, the app is deliberately **stopped and held** rather than started — a running
|
||||||
|
app on a half-written database lets you type into it and makes the damage permanent — and you are
|
||||||
|
told to contact us. That was your ruling this morning. **Two things also stopped:** the error no
|
||||||
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
||||||
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
||||||
|
off-site permanently.
|
||||||
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
||||||
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
||||||
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
||||||
|
|||||||
@@ -352,6 +352,36 @@ above stands exactly as written: the secondary unit mirror is still read by noth
|
|||||||
where the primary drive is lost — and in exactly that case the primary unit is gone while this
|
where the primary drive is lost — and in exactly that case the primary unit is gone while this
|
||||||
mirror survives on the second drive, unreachable by any customer action.
|
mirror survives on the second drive, unreachable by any customer action.
|
||||||
|
|
||||||
|
**[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.**
|
||||||
|
Recorded here rather than only in a closed register row, because a decision that survives only inside
|
||||||
|
a closed work item is a decision nobody will find.
|
||||||
|
|
||||||
|
1. **Replay.** The snapshot's `.sql` is imported into the app's database, with only the DB service up
|
||||||
|
(R-47). A pre-restore copy of the LIVE database was taken first and is on disk (R-43); the restore
|
||||||
|
refuses outright if it could not be taken.
|
||||||
|
2. **Rollback (controller v0.220.0, R-379).** If the replay fails, the product re-applies that undo
|
||||||
|
copy itself. **The whole set for this run** — an app with two databases gets two undo files, and
|
||||||
|
restoring only the first would leave the other half-written — matched on **the run's own stamp**,
|
||||||
|
never on the `pre-restore-` prefix, because several runs' copies coexist in the same directory. It
|
||||||
|
runs with the DB service still up and before any restart, so the app never observes the half
|
||||||
|
state, and into a **re-discovered** container: the DB-only start re-creates it, so the id captured
|
||||||
|
at dump time is dead by rollback time (v0.220.1, found by the first live run). The app then starts
|
||||||
|
and the customer is told **both** that the restore failed and that their data is as it was.
|
||||||
|
3. **Hold (v0.220.0, operator ruling 2026-08-22).** If the rollback ALSO fails, the app is **held
|
||||||
|
stopped**, not started. A running app on a half-written database lets the customer type into it and
|
||||||
|
makes the damage permanent. The hold is persisted, every start path refuses it with a reason and a
|
||||||
|
route, the app-stop marker is ended so nothing auto-restarts it at the next boot, and the app reads
|
||||||
|
**red** rather than green. An operator clears it with `--clear-restore-hold`, which requires a
|
||||||
|
controller restart.
|
||||||
|
|
||||||
|
**Why a rollback and not an engine flag.** Postgres gained `--single-transaction` in the same release
|
||||||
|
and that does make its replay all-or-nothing — but **MariaDB's DDL is not transactional**, so a
|
||||||
|
partial apply there is unavoidable at the engine. Measured 2026-08-22: the same truncated dump left
|
||||||
|
Postgres emptied and crash-looping, and left MariaDB with its user data intact, its schema-version
|
||||||
|
table wiped to zero rows, and the app reporting `health=healthy, running=true, restarts=0`. The flag
|
||||||
|
is a belt; the rollback is the fix. **R-379, R-380.** Evidence:
|
||||||
|
`audits/DRILL-r379-rollback-2026-08-22/`.
|
||||||
|
|
||||||
**[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture
|
**[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture
|
||||||
destination.** The drive if the app declares one (`HDD_PATH`), the system data path otherwise —
|
destination.** The drive if the app declares one (`HDD_PATH`), the system data path otherwise —
|
||||||
`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller
|
`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller
|
||||||
|
|||||||
@@ -0,0 +1,114 @@
|
|||||||
|
# DRILL — R-379/R-380: the rollback, proven live (2026-08-22)
|
||||||
|
|
||||||
|
**Subject:** `demo-hp` (Tier 0), guest 9201. **Controller v0.220.0 → v0.220.1 → v0.220.2** during the
|
||||||
|
walk — the walk itself found two defects and both were fixed and re-proven.
|
||||||
|
**Method:** endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex.
|
||||||
|
**Subjects were FOUND, not rebuilt:** `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), left
|
||||||
|
running with their planted data by the 2026-08-22 R-356b drill.
|
||||||
|
|
||||||
|
## The ladder, proven rung by rung
|
||||||
|
|
||||||
|
| step | subject | result |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | `docmost` (Postgres) | replay forced to fail → **rollback succeeded** → app healthy → data byte-identical |
|
||||||
|
| 2 | `bookstack` (MariaDB) | same → **`migrations` back at 102 rows**, the exact cell R-380 was measured in |
|
||||||
|
| 3 | `docmost` | both failed → **app held, not started**; every start path refused; survived a controller restart; cleared via the operator route |
|
||||||
|
| 4 | `docmost` | clean restore → **no rollback, no hold**, proven with a working positive control |
|
||||||
|
| 5 | `/backups/apps` | no phantom rows; the undo cap holds at 3 per app |
|
||||||
|
|
||||||
|
**Step 1 data check.** Pre-restore: 4 pages, `ALTERED-VALUE-B`, titles sha256
|
||||||
|
`8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19`. Post-restore: **identical on all
|
||||||
|
three**.
|
||||||
|
|
||||||
|
**Step 2 data check.** Pre: entities 1, users 2, migrations 102, book name hex
|
||||||
|
`C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662`, disc `ORIGINAL-VALUE-A`.
|
||||||
|
Post: **identical on all five**.
|
||||||
|
|
||||||
|
## The customer message, verbatim
|
||||||
|
|
||||||
|
**257 bytes** (was 407 on this path, of which ~250 were engine output):
|
||||||
|
|
||||||
|
> A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen — az adataid
|
||||||
|
> visszakerültek a visszaállítás előtti állapotba, az alkalmazás fut tovább. Ha újra megpróbálnád,
|
||||||
|
> előbb vedd fel velünk a kapcsolatot
|
||||||
|
|
||||||
|
Hex in `10-step1-message.txt`. **Judged, not just recorded:** it states the failure AND the recovery.
|
||||||
|
A message reporting only the failure would leave a customer believing their data was gone when it is
|
||||||
|
not. Checked for `ERROR:`, `LINE 1:`, `COPY public.`, `exit status`, `psql`, `INSERT INTO` — **all
|
||||||
|
absent**.
|
||||||
|
|
||||||
|
The held-app message (Scenario C), **verbatim**:
|
||||||
|
|
||||||
|
> a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem
|
||||||
|
> sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább.
|
||||||
|
> Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-…-docmost-postgres.sql
|
||||||
|
|
||||||
|
## Scenario C in plain words
|
||||||
|
|
||||||
|
**What a customer sees:** the app is stopped and stays stopped. Pressing start returns, in Hungarian,
|
||||||
|
that the restore broke, the previous state could not be put back, the app is deliberately stopped so
|
||||||
|
the data cannot be damaged further, and to contact us. The app's row is **red**, not green.
|
||||||
|
|
||||||
|
**What an operator does:** `docker exec felhom-controller /usr/local/bin/felhom-controller
|
||||||
|
--restore-holds` lists the app with both errors. After checking the data (the undo copies are in the
|
||||||
|
app's unit `db-dumps/`), `--clear-restore-hold <app>` clears it — **and then the controller must be
|
||||||
|
restarted**, which the command now says.
|
||||||
|
|
||||||
|
## TWO DEFECTS THE WALK FOUND IN THE FIX ITSELF
|
||||||
|
|
||||||
|
**1. The rollback used a dead container (v0.220.1).** `writeSafetyDump` captures its `DiscoveredDB`
|
||||||
|
before the stop; the DB-only start re-creates the container. Measured: `docmost-postgres` captured as
|
||||||
|
`9adbc14f9af6` at 16:05:44, re-created as `309795897b82` at 16:05:47, rollback's `docker exec` against
|
||||||
|
the dead id sat in `waitDBReady` for 30 s. **So the app was HELD for an infrastructure reason while
|
||||||
|
its data was perfectly recoverable** — the hold behaved correctly on a case that should never have
|
||||||
|
reached it. **No unit test could see this: they all inject the import seam and never look at container
|
||||||
|
identity.** Fixed by re-discovering and matching on `{stack, engine}`; a new test asserts the identity
|
||||||
|
handed to the import, and its red-proof convicts.
|
||||||
|
|
||||||
|
**2. The operator route did not take effect (v0.220.2).** `--clear-restore-hold` runs as a second
|
||||||
|
process: it cleared `settings.json` correctly and the running controller went on refusing, because it
|
||||||
|
holds its own in-memory settings. Found by using the route, not by reading it. The command now prints
|
||||||
|
the restart it needs. The lost-update window between the two processes is recorded rather than hidden.
|
||||||
|
|
||||||
|
## Red-proofs — eight planned, one PASSED
|
||||||
|
|
||||||
|
| # | mutation | observed |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | remove the rollback entirely | 3 tests failed: "the undo copy must be re-applied exactly once, got 0" |
|
||||||
|
| 2 | roll back only the first database | "BOTH databases must be rolled back, got 1" |
|
||||||
|
| 3 | start the app anyway after a failed rollback | "the app was STARTED onto a half-written database" |
|
||||||
|
| 4 | move the hold check below the driveless early return | "the hold check (line 1822) is AFTER the driveless early return (line 1820)" |
|
||||||
|
| 5 | paste engine stderr back into the customer error | **PASSED — reported, not omitted.** See below |
|
||||||
|
| 6 | delete the operator log line | "the engine stderr is no longer logged either" |
|
||||||
|
| 7 | revert the undo naming | "must resolve to the app it belongs to, not a phantom; got `pre-restore-…-docmost`" |
|
||||||
|
| 8 | prune keeps the oldest | "the newest 3 must survive; 20260401T000000Z is gone" |
|
||||||
|
| 9 | use the captured container id | "the rollback used the CAPTURED container id 9adbc14f9af6" |
|
||||||
|
|
||||||
|
**Red-proof 5 passed and that is a finding.** The R-381 behavioural test injects at `m.importDBDump`,
|
||||||
|
i.e. BELOW `ImportDump`, so re-adding the stderr inside `ImportDump` could not fail it — the test was
|
||||||
|
hollow for the layer the leak lives in. A guard was added at that layer (an AST assertion that the
|
||||||
|
returned error does not carry the captured stderr, plus a positive control that it is still logged),
|
||||||
|
and the mutation then convicted.
|
||||||
|
|
||||||
|
## What did NOT reproduce
|
||||||
|
|
||||||
|
**The undo copies rendering as apps on the customer's backup page.** The live page was read BEFORE any
|
||||||
|
change and contained **zero** `pre-restore` strings while four such files sat on disk — with a
|
||||||
|
positive control showing 8 real app rows and a database section. `buildAppBackupRows` iterates
|
||||||
|
DEPLOYED apps and only reads the derived-name map by key, so a phantom name becomes a KEY and never a
|
||||||
|
ROW. The phantom key was real and is fixed; the visible row was not. Their visibility is unchanged and
|
||||||
|
remains deliberate.
|
||||||
|
|
||||||
|
## Teardown — three layers
|
||||||
|
|
||||||
|
1. **Nothing was provisioned.** `docmost` and `bookstack` were already on the box. Both are healthy at
|
||||||
|
the end, with their data byte-identical to the start.
|
||||||
|
2. **No storage was added.** The undo copies are now capped at **3 per app** (was 4 and 2 growing),
|
||||||
|
and the cap was observed firing live: *"pruned old undo copy … (keeping the newest 3)"*.
|
||||||
|
3. **No hub-side record was created.** No customer, no appliance. Nothing to dispose of.
|
||||||
|
|
||||||
|
## Deliberately left open
|
||||||
|
|
||||||
|
R-102 (the Tier-2 unit mirror read by nothing), R-359 (no readability check on the off-site store),
|
||||||
|
R-361 (the app's own dump vanishing from the unit after a reconstitution — reproduced again here and
|
||||||
|
still not fixed). Separate rows, untouched.
|
||||||
+14
@@ -0,0 +1,14 @@
|
|||||||
|
== docmost PRE-RESTORE state
|
||||||
|
pages : 4
|
||||||
|
users : 1
|
||||||
|
disc : ERROR: relation "felhom_r356b_disc" does not exist
|
||||||
|
LINE 1: SELECT marker FROM felhom_r356b_disc
|
||||||
|
^
|
||||||
|
titles (sha256 of the sorted set):
|
||||||
|
8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19 -
|
||||||
|
per-title hex:
|
||||||
|
52333536422d504147452d312d73656e74696e656c
|
||||||
|
52333536422d504147452d322d73656e74696e656c
|
||||||
|
5c753030633172765c75303065647a745c7530313731725c753031353120745c75303066636b5c753030663672665c7530306661725c7530306633675c753030653970205c7532303134205233353662206472696c6c
|
||||||
|
c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a970203220e28094205233353662
|
||||||
|
ALTERED-VALUE-B
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
size before: 141363
|
||||||
|
size after : 62000
|
||||||
|
tail: public; Owner: ---COPY public.felhom_r356b_discriminator
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
2026/08/22 16:05:41 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/reconstitute
|
||||||
|
2026/08/22 16:05:41 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/reconstitute from 172.18.0.4:54196
|
||||||
|
2026/08/22 16:05:44 offbox_reconstitute.go:200: [INFO] [offbox] docmost: pre-restore safety dump written → pre-restore-20260822T160544Z-docmost-postgres.sql (138.0 KB)
|
||||||
|
2026/08/22 16:05:48 offbox_reconstitute.go:642: [ERROR] [offbox] docmost: database replay failed, rolling back to the pre-restore state: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3
|
||||||
|
2026/08/22 16:05:48 offbox_reconstitute.go:349: [INFO] [offbox] docmost: rolling back to the pre-restore state from pre-restore-20260822T160544Z-docmost-postgres.sql
|
||||||
|
2026/08/22 16:06:19 offbox_reconstitute.go:647: [ERROR] [offbox] docmost: ROLLBACK ALSO FAILED (a korábbi állapot visszaállítása sikertelen (docmost-postgres): waiting for docmost-postgres (postgres) readiness: timeout after 30s) — holding the app stopped; replay error was: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3
|
||||||
|
2026/08/22 16:06:19 offbox_handlers.go:451: [ERROR] [web] off-box reconstitute docmost (async): a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-20260822T160544Z-docmost-postgres.sql
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
=== operator CLI: --restore-holds
|
||||||
|
docmost held since 2026-08-22T16:06:19Z
|
||||||
|
replay error : importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3
|
||||||
|
rollback err : a korábbi állapot visszaállítása sikertelen (docmost-postgres): waiting for docmost-postgres (postgres) readiness: timeout after 30s
|
||||||
|
|
||||||
|
=== the CUSTOMER's start button (POST /api/stacks/docmost/start):
|
||||||
|
{"ok":false,"error":"a(z) docmost adatainak visszaállítása 2026-08-22 16:06-kor megszakadt, és a korábbi állapotot sem sikerült visszatölteni. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot"}
|
||||||
+14
@@ -0,0 +1,14 @@
|
|||||||
|
=== after a controller RESTART — did anything start the held app?
|
||||||
|
docmost-postgres | Up 2 minutes (healthy)
|
||||||
|
|
||||||
|
=== the boot sweep / Recover lines:
|
||||||
|
2026/08/22 16:07:52 sync.go:371: [DEBUG] [sync] docmost/docker-compose.yml: hash match, skipped
|
||||||
|
2026/08/22 16:07:52 sync.go:371: [DEBUG] [sync] docmost/.felhom.yml: hash match, skipped
|
||||||
|
309795897b82 docmost-postgres docmost postgres:16-alpine
|
||||||
|
2026/08/22 16:07:52 dbdump.go:176: [DEBUG] DiscoverDatabases: found postgres container: docmost-postgres (id=309795897b82)
|
||||||
|
2026/08/22 16:07:52 dbdump.go:196: [DEBUG] DiscoverDatabases: docmost-postgres → stack=docmost, dbUser=docmost, dbName=docmost
|
||||||
|
2026/08/22 16:07:52 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
|
||||||
|
2026/08/22 16:07:52 recovery_unit.go:194: [INFO] [backup] Recovery unit captured for docmost → /mnt/sys_drive/felhom-data/backups/primary/docmost (images=3, secrets-referenced=2, data_keys=0, portable-carried=2/2, withheld=0)
|
||||||
|
2026/08/22 16:07:52 offbox_reconstitute.go:256: [INFO] [backup] docmost: pruned old undo copy pre-restore-20260822T140501Z-docmost-postgres.sql (keeping the newest 3)
|
||||||
|
2026/08/22 16:07:58 logscanner.go:69: [DEBUG] [metrics] logscanner: scanned docmost-postgres: errors=1 warnings=0 issues=1 (took 22ms)
|
||||||
|
2026/08/22 16:08:02 healthprobe.go:162: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
restore hold cleared for docmost — the app may be started again. Check its data first: the undo copies are in its unit's db-dumps dir.
|
||||||
|
{"ok":false,"error":"a(z) docmost adatainak visszaállítása 2026-08-22 16:06-kor megszakadt, és a korábbi állapotot sem sikerült visszatölteni. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot"}
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
4
|
||||||
|
ALTERED-VALUE-B
|
||||||
|
8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19 -
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
2026/08/22 16:23:44 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/reconstitute
|
||||||
|
2026/08/22 16:23:44 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/reconstitute from 172.18.0.4:41900
|
||||||
|
2026/08/22 16:23:47 offbox_reconstitute.go:200: [INFO] [offbox] docmost: pre-restore safety dump written → pre-restore-20260822T162347Z-docmost-postgres.sql (138.0 KB)
|
||||||
|
2026/08/22 16:23:51 offbox_reconstitute.go:690: [ERROR] [offbox] docmost: database replay failed, rolling back to the pre-restore state: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3
|
||||||
|
2026/08/22 16:23:52 offbox_reconstitute.go:394: [DEBUG] [offbox] docmost: docmost-postgres was re-created during the restore (309795897b82 → 48817bfb454a) — rolling back into the live container
|
||||||
|
2026/08/22 16:23:52 offbox_reconstitute.go:397: [INFO] [offbox] docmost: rolling back to the pre-restore state from pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||||
|
2026/08/22 16:23:53 offbox_reconstitute.go:402: [INFO] [offbox] docmost: rollback complete — 1 database(s) returned to the pre-restore state
|
||||||
|
2026/08/22 16:24:04 offbox_handlers.go:451: [ERROR] [web] off-box reconstitute docmost (async): a(z) docmost adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
=== app state:
|
||||||
|
docmost | Up 29 seconds (healthy)
|
||||||
|
docmost-redis | Up 39 seconds (healthy)
|
||||||
|
docmost-postgres | Up 43 seconds (healthy)
|
||||||
|
=== POST-RESTORE data (must equal the pre-restore state exactly):
|
||||||
|
4
|
||||||
|
ALTERED-VALUE-B
|
||||||
|
8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19 -
|
||||||
|
expected: 4 / ALTERED-VALUE-B / 8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19
|
||||||
|
=== was a hold written? (must be none)
|
||||||
|
restore_holds = None
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
=== CUSTOMER MESSAGE, VERBATIM (Postgres, rollback succeeded) ===
|
||||||
|
A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot
|
||||||
|
|
||||||
|
=== UTF-8 hex ===
|
||||||
|
412074656c6a657320766973737a61c3a16c6cc3ad74c3a1732073696b657274656c656e3a2061287a2920646f636d6f7374206164617462c3a17a6973c3a16e616b20766973737a61c3a16c6cc3ad74c3a173612073696b657274656c656e20e2809420617a206164617461696420766973737a616b6572c3bc6c74656b206120766973737a61c3a16c6cc3ad74c3a17320656cc59174746920c3a16c6c61706f7462612c20617a20616c6b616c6d617ac3a1732066757420746f76c3a162622e20486120c3ba6a7261206d65677072c3b362c3a16c6ec3a1642c20656cc591626220766564642066656c2076656cc3bc6e6b2061206b617063736f6c61746f74
|
||||||
|
|
||||||
|
=== byte length: 257
|
||||||
|
|
||||||
|
=== engine-output leak check:
|
||||||
|
ERROR: present: False
|
||||||
|
LINE 1: present: False
|
||||||
|
COPY public. present: False
|
||||||
|
exit status present: False
|
||||||
|
psql present: False
|
||||||
|
INSERT INTO present: False
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
entities : 1
|
||||||
|
users : 2
|
||||||
|
migrations: 102
|
||||||
|
book hex : C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662
|
||||||
|
disc : ORIGINAL-VALUE-A
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
2026/08/22 16:25:56 offbox_reconstitute.go:200: [INFO] [offbox] bookstack: pre-restore safety dump written → pre-restore-20260822T162555Z-bookstack-mariadb.sql (57.4 KB)
|
||||||
|
2026/08/22 16:26:05 offbox_reconstitute.go:690: [ERROR] [offbox] bookstack: database replay failed, rolling back to the pre-restore state: importing mariadb dump for bookstack: mariadb import into bookstack-db failed: exit status 1
|
||||||
|
2026/08/22 16:26:05 offbox_reconstitute.go:394: [DEBUG] [offbox] bookstack: bookstack-db was re-created during the restore (0f251288ec4a → e5895283fa31) — rolling back into the live container
|
||||||
|
2026/08/22 16:26:05 offbox_reconstitute.go:397: [INFO] [offbox] bookstack: rolling back to the pre-restore state from pre-restore-20260822T162555Z-bookstack-mariadb.sql
|
||||||
|
2026/08/22 16:26:06 offbox_reconstitute.go:402: [INFO] [offbox] bookstack: rollback complete — 1 database(s) returned to the pre-restore state
|
||||||
|
2026/08/22 16:26:08 offbox_handlers.go:451: [ERROR] [web] off-box reconstitute bookstack (async): a(z) bookstack adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
entities : 1
|
||||||
|
users : 2
|
||||||
|
migrations: 102
|
||||||
|
book hex : C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662
|
||||||
|
disc : ORIGINAL-VALUE-A
|
||||||
+15
@@ -0,0 +1,15 @@
|
|||||||
|
=== the run:
|
||||||
|
2026/08/22 16:27:37 offbox_reconstitute.go:722: [INFO] [offbox] reconstituted docmost from snapshot 750b7b4d: 0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed, safety dump=pre-restore-20260822T162708Z-docmost-postgres.sql, skewed=false
|
||||||
|
=== ABSENCE CLAIMS — no rollback, no hold:
|
||||||
|
rollback lines in this run: 3
|
||||||
|
restore_holds: None
|
||||||
|
=== This run's window: safety dump 16:27:08 → reconstituted 16:27:37
|
||||||
|
=== every 'rolling back' line today, with timestamps:
|
||||||
|
2026/08/22 16:23:52 offbox_reconstitute.go:397: [INFO] [offbox] docmost: rolling back to the pre-restore state from pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||||
|
2026/08/22 16:26:05 offbox_reconstitute.go:397: [INFO] [offbox] bookstack: rolling back to the pre-restore state from pre-restore-20260822T162555Z-bookstack-mariadb.sql
|
||||||
|
|
||||||
|
=== POSITIVE CONTROL: the grep above finds rollback lines (it printed some).
|
||||||
|
=== NEGATIVE: none of them falls between 16:27:08 and 16:27:37 — this run rolled back nothing.
|
||||||
|
|
||||||
|
=== and the R-382 fix, visible on the same line:
|
||||||
|
2026/08/22 16:27:37 offbox_reconstitute.go:722: [INFO] [offbox] reconstituted docmost from snapshot 750b7b4d: 0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed, safety dump=pre-restore-20260822T162708Z-docmost-postgres.sql, skewed=false
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
=== LIVE /backups/apps on 0.220.2
|
||||||
|
'pre-restore' occurrences on the page : 0
|
||||||
|
real app rows (positive control) : bookstack docmost kimai privatebin
|
||||||
|
undo copies on disk right now : 6
|
||||||
|
|
||||||
|
So: undo copies exist, the page renders real apps, and no phantom row appears.
|
||||||
|
The reported symptom did NOT reproduce on v0.219.0 either — it was read on the live page
|
||||||
|
BEFORE any change was made. What was real was the phantom map KEY, now fixed.
|
||||||
|
|
||||||
|
=== the prune cap in force (max 3 per app):
|
||||||
|
docmost 3 undo copies
|
||||||
|
bookstack 3 undo copies
|
||||||
@@ -26,6 +26,10 @@
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
| **R-379** | **The pre-restore undo copy was valid, was named to the customer, and no product action could apply it.** Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/`. **Reasoning kept:** *R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken.* **The undo set is matched on THE RUN'S OWN STAMP, never on the `pre-restore-` prefix** (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) **and never just the first file** (a two-database app would have had one restored and the other left half-written). **The rollback RE-DISCOVERS the container** — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of `waitDBReady` against a dead id while its data was recoverable. **No unit test saw that: they all inject the import seam and never look at container identity.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
| **R-380** | **A failed MariaDB replay left a partially-applied database behind an app reporting `health=healthy`.** Shipped in controller v0.220.0. Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt`. **Reasoning kept:** **no engine flag closes this** — `--single-transaction` was added to the Postgres import and does make it all-or-nothing, but **MariaDB's DDL is not transactional**, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: `bookstack`'s `migrations` table back at **102 rows**, the exact cell the defect was measured in. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
| **R-381** | **The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface.** Shipped in controller v0.220.0. **Reasoning kept:** the full engine text now goes to the operator log, **which never had it before — the diagnostic was ADDED, not removed**. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an `INSERT INTO migrations VALUES (…)` listing); now 257 bytes with no engine tokens. **A red-proof for this PASSED and the test was hollow**: it injected below `ImportDump`, so a leak reintroduced inside `ImportDump` could not fail it. The guard now sits at that layer. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
| **R-382** | **The reconstitution's summary log line omitted the volume count it already held.** Shipped in controller v0.220.0. Proven live: `0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed`. | **CLOSED — SHIPPED** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
| **R-356** | **The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no.** Shipped in controller v0.219.0. Evidence: `audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. **Reasoning kept:** *the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (`Manager.GetAppDrivePath`, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong.* **An app with no drive is not misconfigured** — `01-topology-and-trust.md` §8 carries the `[DESIGN]` marker; between 19 and 22 August that design was called a defect four times. **Deployment is asked of `ListDeployedStacks()` and FAILS CLOSED on a nil provider:** "cannot tell" must not become "go ahead" when the caller's next act is a write. **Two different failures get two different sentences** — installed-but-no-resolvable-data-root has its own refusal and its own route; widening `nincs telepítve` to cover it would send a customer to reinstall a running app and hide the real fault. **Measured, and load-bearing: 53 templates, 13 `needs_hdd: true`, 40 `false`** (catalogue @ `459766cb1639`). **The capture side's raw `GetStackHDDPath` is FENCED and was not changed** — capture resolves an app's declared `userdata`/`import` file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.219.0, 2026-08-22; `privatebin` on `demo-hp`: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: `git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` |
|
| **R-356** | **The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no.** Shipped in controller v0.219.0. Evidence: `audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. **Reasoning kept:** *the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (`Manager.GetAppDrivePath`, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong.* **An app with no drive is not misconfigured** — `01-topology-and-trust.md` §8 carries the `[DESIGN]` marker; between 19 and 22 August that design was called a defect four times. **Deployment is asked of `ListDeployedStacks()` and FAILS CLOSED on a nil provider:** "cannot tell" must not become "go ahead" when the caller's next act is a write. **Two different failures get two different sentences** — installed-but-no-resolvable-data-root has its own refusal and its own route; widening `nincs telepítve` to cover it would send a customer to reinstall a running app and hide the real fault. **Measured, and load-bearing: 53 templates, 13 `needs_hdd: true`, 40 `false`** (catalogue @ `459766cb1639`). **The capture side's raw `GetStackHDDPath` is FENCED and was not changed** — capture resolves an app's declared `userdata`/`import` file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.219.0, 2026-08-22; `privatebin` on `demo-hp`: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: `git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
| **R-216** | **A correct recovery code was reported to the customer as wrong.** Shipped in 0.120.0, v0.125.0. | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
| **R-216** | **A correct recovery code was reported to the customer as wrong.** Shipped in 0.120.0, v0.125.0. | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** Shipped in v0.203.0. Evidence: `documentation/tests/part4-rewalk-2026-08-06/journal.md`. | **CLOSED 2026-08-06 — controller v0.203.0, proven live.** *(State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.)* **The over-claim, as it stood: the fix covered the DECLARATION half only.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** Shipped in v0.203.0. Evidence: `documentation/tests/part4-rewalk-2026-08-06/journal.md`. | **CLOSED 2026-08-06 — controller v0.203.0, proven live.** *(State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.)* **The over-claim, as it stood: the fix covered the DECLARATION half only.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
|||||||
@@ -136,10 +136,6 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
|||||||
|
|
||||||
| ID | What | State |
|
| ID | What | State |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| **R-379** | **The pre-restore undo copy is taken, is valid, is named to the customer — and NOTHING IN THE PRODUCT CAN APPLY IT.** When an off-site DB replay fails, `reimportDBDumpsFrom` returns and the refusal names the safety dump by filename (`offbox_reconstitute.go:436-441`). That filename appears ONLY inside the error string: there is no button, no list entry, no route. `preRestoreDumpPrefix` ("pre-restore-") is deliberately SKIPPED at three sites so these files are never offered as a restore source — `internal/backup/restore_unit.go:125`, `internal/backup/offbox_reconstitute.go:489`, `internal/backup/offbox_reconstitute.go:539`. **PROVEN LIVE 2026-08-22 on `demo-hp`, both engines.** Postgres (`docmost`): after a truncated dump the live database held 43 tables and **0 rows** in `pages`, `users` and `spaces`, and the app crash-looped. The undo copy (141 363 B, 43 COPY blocks, 4 page rows, 1 user row, the accented title present) was applied BY HAND and restored the exact prior state. MariaDB (`bookstack`): same, `migrations` 0 -> 102 rows. **So the data is recoverable — by us, by hand, over a support conversation. The customer has a filename.** | **OPEN — HIGH** | — | Offer the undo copy as a restore source on the app's restore page when one exists, or state in the message that recovery needs support and how to ask. The skip at the three sites is correct for *normal* listing — the gap is that there is no deliberate second surface. **Do NOT widen the three skips**: they exist so a safety dump is never mistaken for the app's own backup (that confusion is R-361's neighbourhood). | CC |
|
|
||||||
| **R-380** | **A failed MariaDB replay leaves a PARTIALLY APPLIED database behind an app that reports HEALTHY — Postgres fails visibly, MariaDB does not.** `ImportDump` gives the Postgres branch `-v ON_ERROR_STOP=1`; the MariaDB branch is a plain `mariadb -u root -p<pw> <db>` with no equivalent (`internal/appbackup/dbdump.go:670-692`). Both DO surface the failure — H3's predicted 'quiet success' did NOT occur — but the STATE they leave differs, and that is the defect. **Measured 2026-08-22 on `demo-hp` with the same truncation on both engines.** Postgres: everything emptied, app crash-loops, `Restarting (1)` — visibly broken. MariaDB: the dump's DROP/CREATE/INSERT runs table by table, so tables it reached are rebuilt, tables it never reached keep their ORIGINAL data, and the table it died inside is left EMPTY. Result on `bookstack`: `entities` 1 (intact), `users` 2 (intact), **`migrations` 0 rows (wiped)** — the schema-version ledger — while `docker inspect` reported **`health=healthy running=true restarts=0`** and the app served HTTP. An empty `migrations` table means BookStack believes no migration has ever run; the next upgrade would re-run all 102 against an existing schema. **Nothing signals ongoing damage.** | **OPEN — HIGH** | — | Make a failed replay leave a KNOWN state rather than a partial one: wrap the MariaDB import so a failure is atomic, or re-apply the undo copy automatically on import failure (which needs R-379 first), or at minimum mark the app unhealthy so the dashboard stops saying it is fine. **The MariaDB client's default IS to abort on error — that was measured, not assumed — so this is not a missing flag; it is the absence of a transaction boundary.** | CC |
|
|
||||||
| **R-381** | **The restore-failure message pastes raw database-engine stderr — including the customer's own database rows — into a Hungarian customer-facing surface.** `ImportDump` truncates stderr to 300 chars and wraps it verbatim (`internal/appbackup/dbdump.go:700-706`); `offbox_reconstitute.go:436-441` wraps that again; the flash renders it in an `alert alert-error` block. **Measured verbatim 2026-08-22.** Postgres, **407 bytes**, of which ~250 are untranslated English psql output with a caret diagram and `exit status 3`. MariaDB, **615 bytes**, whose middle is an `INSERT INTO \`migrations\` VALUES (1,'2014_10_12_000000_create_users_table',1),(2,...` listing — i.e. **actual table contents, HTML-escaped, shown to the customer**. On a real app that statement could be any row the dump died inside. | **OPEN — MEDIUM** | — | Keep the engine text in the operator log where it belongs; give the customer the reason, the undo copy and the route. Same class as **R-79** (English on customer surfaces) and **R-257** (internal state names in customer copy), but a distinct producer and with a content-disclosure dimension neither has: this one can print rows. | CC |
|
|
||||||
| **R-382** | **The reconstitution's summary log line omits the volume count it already computed.** `offbox_reconstitute.go:452` logs `%d file(s) placed, %d DB dump(s) replayed, safety dump=%s, skewed=%v` — `res.VolumesReplayed` is set at line 412 and never printed. Measured 2026-08-22: `docmost` logged `0 file(s) placed, 1 DB dump(s) replayed` on a run that replayed **3** volumes including the entire 52 MB Postgres data directory; `bookstack` logged the same shape on a run that replayed 2 including a 161 MB one. The customer-facing flash DOES name the volumes („0 fájl és 1 adatkötet visszaállítva") — so the operator log is less informative than the customer message. This directly obstructed answering the 2026-08-22 drill's Q2 from the log and forced a planted discriminator instead. | **OPEN — LOW** | — | Add `%d volume(s) replayed` to the line. One format string. | CC |
|
|
||||||
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
||||||
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
|
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
|
||||||
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |
|
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |
|
||||||
|
|||||||
@@ -0,0 +1,65 @@
|
|||||||
|
# Golden bake — 0.220.2 (2026-08-22)
|
||||||
|
|
||||||
|
Baked in the drill VM on DooPlex per `documentation/runbooks/RUNBOOK-manual-build.md` §4.0/§4.1,
|
||||||
|
carrying controller **v0.220.2** (R-379/R-380/R-381/R-382 — the rollback, the hold, and the two
|
||||||
|
defects the live walk found in them).
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| `GOLDEN_VERSION` | **0.220.2** |
|
||||||
|
| `GOLDEN_SHA256` | **cb439418c7005ce01bcb8126bb6385688c2f408c3c4649f6740001bc2c864ed5** |
|
||||||
|
| package | `https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.220.2/golden.tar.zst` |
|
||||||
|
| size | 657 271 965 B |
|
||||||
|
| controller image | `gitea.dooplex.hu/admin/felhom-controller:0.220.2` |
|
||||||
|
| template | `debian-13-standard_13.6-1_amd64.tar.zst` (**listed fresh**, checksum verified on download) |
|
||||||
|
| `build-golden.sh` | v3.0.0, from `felhom-agent` @ `40d857b52711` |
|
||||||
|
| `MinAgent` | **0.129.0** (from the controller CHANGELOG header — unchanged) |
|
||||||
|
|
||||||
|
## Why this bake happened in this session
|
||||||
|
|
||||||
|
The `golden-currency` gate refused the docs push: controller 0.220.2 was released with no golden
|
||||||
|
carrying it. **That block is not circular** — a golden needs the controller image, which was already
|
||||||
|
built and pushed, not the docs commit. So the gate was satisfied by doing the work it asked for,
|
||||||
|
rather than bypassed with `--no-verify`.
|
||||||
|
|
||||||
|
## Pass markers — each checked, with the negative controls
|
||||||
|
|
||||||
|
```
|
||||||
|
docker OK (overlay2 : 1 -> " docker OK (overlay2; data-root /var/lib/docker)"
|
||||||
|
including mount point : 2 -> rootfs ('/') and mp0 ('/var/lib/felhom') [there is no mp1]
|
||||||
|
upload OK (HTTP 201) : 1 -> pre-delete returned HTTP 404 (404/204 expected)
|
||||||
|
excluding : 0 <- negative control
|
||||||
|
FATAL : 0 <- negative control
|
||||||
|
```
|
||||||
|
|
||||||
|
**The 404 pre-gate was controlled before it was believed:** the same URL shape for **0.219.0 returned
|
||||||
|
HTTP 200** in the same minute, so the 404 on 0.220.2 means absent, not a wrong URL.
|
||||||
|
|
||||||
|
## Verified by ROUND TRIP
|
||||||
|
|
||||||
|
The published object was downloaded again — **HTTP 200, 657 271 965 bytes** — and its sha256
|
||||||
|
recomputed: `cb439418…` on both sides. What a machine receives is byte-identical to what was baked.
|
||||||
|
|
||||||
|
## Token hygiene
|
||||||
|
|
||||||
|
Copied **file → file** and read by a runner script inside the VM.
|
||||||
|
`systemctl show golden-bake -p Environment -p ExecStart | grep -c -F "$(cat /root/.gitea-token)"` → **0**,
|
||||||
|
and that grep was **proved able to convict first** (planted copy → **1**, copy shredded).
|
||||||
|
The same control was run on the **committed** `bake.log`: **0**, positive control **1**.
|
||||||
|
|
||||||
|
## Teardown
|
||||||
|
|
||||||
|
`pct destroy 9100 --purge`; token, runner, build script and in-VM log `shred -u`'d **after** `bake.log`
|
||||||
|
was copied out to this directory — all four confirmed absent; `poweroff`; waited for qemu to exit;
|
||||||
|
`qemu-img snapshot -a virgin`.
|
||||||
|
|
||||||
|
## NOT vouched
|
||||||
|
|
||||||
|
The hub's Day-0 artifact manifest was **not** changed and the floor was **not** raised. Both are the
|
||||||
|
operator's, and it is a **three-field** save:
|
||||||
|
|
||||||
|
| field | value |
|
||||||
|
|---|---|
|
||||||
|
| `golden_version` | **0.220.2** |
|
||||||
|
| `agent_version` | **0.130.0** |
|
||||||
|
| `min_agent` | **0.129.0** |
|
||||||
@@ -0,0 +1,325 @@
|
|||||||
|
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.220.2
|
||||||
|
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
|
||||||
|
Logical volume "vm-9100-disk-0" created.
|
||||||
|
Logical volume pve/vm-9100-disk-0 changed.
|
||||||
|
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||||
|
Filesystem UUID: 0d64b7d6-7aaa-48f0-af9b-affbb4f803aa
|
||||||
|
Superblock backups stored on blocks:
|
||||||
|
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||||
|
4096000, 7962624
|
||||||
|
Logical volume "vm-9100-disk-1" created.
|
||||||
|
Logical volume pve/vm-9100-disk-1 changed.
|
||||||
|
Creating filesystem with 6291456 4k blocks and 1572864 inodes
|
||||||
|
Filesystem UUID: ac639ff8-e9cf-4b7c-9938-e379acec78e9
|
||||||
|
Superblock backups stored on blocks:
|
||||||
|
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||||
|
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
|
||||||
|
Total bytes read: 553512960 (528MiB, 97MiB/s)
|
||||||
|
Detected container architecture: amd64
|
||||||
|
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
|
||||||
|
done: SHA256:hwBeMD4w5PElp4dDVRT4eTbhhajRJBa3/Q1LLn9bT0g root@felhom-golden
|
||||||
|
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
|
||||||
|
done: SHA256:QvTmAFW3Had7tj3Uoz+LoCBF/bLpfMjkRDmI9yNFeU8 root@felhom-golden
|
||||||
|
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
|
||||||
|
done: SHA256:/dw8Ehg69wnYlNaGkDVtHV1qocB8JM4280GeTLkj6g4 root@felhom-golden
|
||||||
|
[golden] starting + installing Docker (official repo, trixie channel) …
|
||||||
|
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||||
|
perl: warning: Setting locale failed.
|
||||||
|
perl: warning: Please check that your locale settings:
|
||||||
|
LANGUAGE = (unset),
|
||||||
|
LC_ALL = (unset),
|
||||||
|
LC_CTYPE = (unset),
|
||||||
|
LC_NUMERIC = (unset),
|
||||||
|
LC_COLLATE = (unset),
|
||||||
|
LC_TIME = (unset),
|
||||||
|
LC_MESSAGES = (unset),
|
||||||
|
LC_MONETARY = (unset),
|
||||||
|
LC_ADDRESS = (unset),
|
||||||
|
LC_IDENTIFICATION = (unset),
|
||||||
|
LC_MEASUREMENT = (unset),
|
||||||
|
LC_PAPER = (unset),
|
||||||
|
LC_TELEPHONE = (unset),
|
||||||
|
LC_NAME = (unset),
|
||||||
|
LANG = "en_US.UTF-8"
|
||||||
|
are supported and installed on your system.
|
||||||
|
perl: warning: Falling back to the standard locale ("C").
|
||||||
|
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||||
|
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||||
|
perl: warning: Setting locale failed.
|
||||||
|
perl: warning: Please check that your locale settings:
|
||||||
|
LANGUAGE = (unset),
|
||||||
|
LC_ALL = (unset),
|
||||||
|
LC_CTYPE = (unset),
|
||||||
|
LC_NUMERIC = (unset),
|
||||||
|
LC_COLLATE = (unset),
|
||||||
|
LC_TIME = (unset),
|
||||||
|
LC_MESSAGES = (unset),
|
||||||
|
LC_MONETARY = (unset),
|
||||||
|
LC_ADDRESS = (unset),
|
||||||
|
LC_IDENTIFICATION = (unset),
|
||||||
|
LC_MEASUREMENT = (unset),
|
||||||
|
LC_PAPER = (unset),
|
||||||
|
LC_TELEPHONE = (unset),
|
||||||
|
LC_NAME = (unset),
|
||||||
|
LANG = "en_US.UTF-8"
|
||||||
|
are supported and installed on your system.
|
||||||
|
perl: warning: Falling back to the standard locale ("C").
|
||||||
|
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||||
|
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
|
||||||
|
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
|
||||||
|
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
|
||||||
|
Unable to find image 'hello-world:latest' locally
|
||||||
|
latest: Pulling from library/hello-world
|
||||||
|
4f55086f7dd0: Pulling fs layer
|
||||||
|
4f55086f7dd0: Verifying Checksum
|
||||||
|
4f55086f7dd0: Download complete
|
||||||
|
4f55086f7dd0: Pull complete
|
||||||
|
Digest: sha256:5dd0d3e6e255913fc30f90b9f2b1d359cc2cbdb48090cc4b65f1676e203243cc
|
||||||
|
Status: Downloaded newer image for hello-world:latest
|
||||||
|
docker OK (overlay2; data-root /var/lib/docker)
|
||||||
|
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
|
||||||
|
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
|
||||||
|
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
|
||||||
|
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.220.2 (no registry cred at deploy) …
|
||||||
|
|
||||||
|
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
|
||||||
|
Configure a credential helper to remove this warning. See
|
||||||
|
https://docs.docker.com/go/credential-store/
|
||||||
|
|
||||||
|
0.220.2: Pulling from admin/felhom-controller
|
||||||
|
039e6f9f9752: Pulling fs layer
|
||||||
|
0094c3ac0914: Pulling fs layer
|
||||||
|
deca1dac7403: Pulling fs layer
|
||||||
|
11c19a33d1b8: Pulling fs layer
|
||||||
|
4f3e54e4eec5: Pulling fs layer
|
||||||
|
37c5f038fab8: Pulling fs layer
|
||||||
|
11c19a33d1b8: Waiting
|
||||||
|
4f3e54e4eec5: Waiting
|
||||||
|
37c5f038fab8: Waiting
|
||||||
|
deca1dac7403: Verifying Checksum
|
||||||
|
deca1dac7403: Download complete
|
||||||
|
11c19a33d1b8: Verifying Checksum
|
||||||
|
11c19a33d1b8: Download complete
|
||||||
|
4f3e54e4eec5: Verifying Checksum
|
||||||
|
4f3e54e4eec5: Download complete
|
||||||
|
37c5f038fab8: Verifying Checksum
|
||||||
|
37c5f038fab8: Download complete
|
||||||
|
0094c3ac0914: Verifying Checksum
|
||||||
|
0094c3ac0914: Download complete
|
||||||
|
039e6f9f9752: Verifying Checksum
|
||||||
|
039e6f9f9752: Download complete
|
||||||
|
039e6f9f9752: Pull complete
|
||||||
|
0094c3ac0914: Pull complete
|
||||||
|
deca1dac7403: Pull complete
|
||||||
|
11c19a33d1b8: Pull complete
|
||||||
|
4f3e54e4eec5: Pull complete
|
||||||
|
37c5f038fab8: Pull complete
|
||||||
|
Digest: sha256:3ad3862fff66746539781b686a78915525ab2438ea1b3c9fa4ee978b75ee7c69
|
||||||
|
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.220.2
|
||||||
|
gitea.dooplex.hu/admin/felhom-controller:0.220.2
|
||||||
|
[golden] asking the controller which infra images it manages …
|
||||||
|
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
|
||||||
|
v3.6.7: Pulling from library/traefik
|
||||||
|
589002ba0eae: Pulling fs layer
|
||||||
|
ef63511ea6cc: Pulling fs layer
|
||||||
|
0738e5cb835e: Pulling fs layer
|
||||||
|
3e6813f70c64: Pulling fs layer
|
||||||
|
3e6813f70c64: Waiting
|
||||||
|
589002ba0eae: Verifying Checksum
|
||||||
|
589002ba0eae: Download complete
|
||||||
|
ef63511ea6cc: Verifying Checksum
|
||||||
|
ef63511ea6cc: Download complete
|
||||||
|
3e6813f70c64: Download complete
|
||||||
|
0738e5cb835e: Verifying Checksum
|
||||||
|
0738e5cb835e: Download complete
|
||||||
|
589002ba0eae: Pull complete
|
||||||
|
ef63511ea6cc: Pull complete
|
||||||
|
0738e5cb835e: Pull complete
|
||||||
|
3e6813f70c64: Pull complete
|
||||||
|
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
|
||||||
|
Status: Downloaded newer image for traefik:v3.6.7
|
||||||
|
docker.io/library/traefik:v3.6.7
|
||||||
|
2026.6.0: Pulling from cloudflare/cloudflared
|
||||||
|
47de5dd0b812: Pulling fs layer
|
||||||
|
c172f21841df: Pulling fs layer
|
||||||
|
99515e7b4d35: Pulling fs layer
|
||||||
|
99ba982a9142: Pulling fs layer
|
||||||
|
d6b1b89eccac: Pulling fs layer
|
||||||
|
2780920e5dbf: Pulling fs layer
|
||||||
|
7c12895b777b: Pulling fs layer
|
||||||
|
3214acf345c0: Pulling fs layer
|
||||||
|
52630fc75a18: Pulling fs layer
|
||||||
|
dd64bf2dd177: Pulling fs layer
|
||||||
|
b839dfae01f6: Pulling fs layer
|
||||||
|
ebddc55facdc: Pulling fs layer
|
||||||
|
bdfd7f7e5bf6: Pulling fs layer
|
||||||
|
2d4d7adf6272: Pulling fs layer
|
||||||
|
40008157d8d2: Pulling fs layer
|
||||||
|
bd8962e29291: Pulling fs layer
|
||||||
|
cac2ae0193cb: Pulling fs layer
|
||||||
|
74d1dac84ecc: Pulling fs layer
|
||||||
|
dd64bf2dd177: Waiting
|
||||||
|
b839dfae01f6: Waiting
|
||||||
|
ebddc55facdc: Waiting
|
||||||
|
bdfd7f7e5bf6: Waiting
|
||||||
|
2d4d7adf6272: Waiting
|
||||||
|
99ba982a9142: Waiting
|
||||||
|
d6b1b89eccac: Waiting
|
||||||
|
7c12895b777b: Waiting
|
||||||
|
3214acf345c0: Waiting
|
||||||
|
52630fc75a18: Waiting
|
||||||
|
2780920e5dbf: Waiting
|
||||||
|
40008157d8d2: Waiting
|
||||||
|
bd8962e29291: Waiting
|
||||||
|
cac2ae0193cb: Waiting
|
||||||
|
74d1dac84ecc: Waiting
|
||||||
|
47de5dd0b812: Download complete
|
||||||
|
99515e7b4d35: Verifying Checksum
|
||||||
|
99515e7b4d35: Download complete
|
||||||
|
c172f21841df: Verifying Checksum
|
||||||
|
c172f21841df: Download complete
|
||||||
|
47de5dd0b812: Pull complete
|
||||||
|
d6b1b89eccac: Verifying Checksum
|
||||||
|
d6b1b89eccac: Download complete
|
||||||
|
2780920e5dbf: Verifying Checksum
|
||||||
|
2780920e5dbf: Download complete
|
||||||
|
99ba982a9142: Verifying Checksum
|
||||||
|
99ba982a9142: Download complete
|
||||||
|
c172f21841df: Pull complete
|
||||||
|
7c12895b777b: Verifying Checksum
|
||||||
|
7c12895b777b: Download complete
|
||||||
|
3214acf345c0: Verifying Checksum
|
||||||
|
3214acf345c0: Download complete
|
||||||
|
52630fc75a18: Download complete
|
||||||
|
dd64bf2dd177: Verifying Checksum
|
||||||
|
dd64bf2dd177: Download complete
|
||||||
|
b839dfae01f6: Verifying Checksum
|
||||||
|
b839dfae01f6: Download complete
|
||||||
|
ebddc55facdc: Verifying Checksum
|
||||||
|
ebddc55facdc: Download complete
|
||||||
|
bdfd7f7e5bf6: Verifying Checksum
|
||||||
|
bdfd7f7e5bf6: Download complete
|
||||||
|
40008157d8d2: Verifying Checksum
|
||||||
|
40008157d8d2: Download complete
|
||||||
|
bd8962e29291: Download complete
|
||||||
|
99515e7b4d35: Pull complete
|
||||||
|
2d4d7adf6272: Verifying Checksum
|
||||||
|
2d4d7adf6272: Download complete
|
||||||
|
cac2ae0193cb: Verifying Checksum
|
||||||
|
cac2ae0193cb: Download complete
|
||||||
|
74d1dac84ecc: Verifying Checksum
|
||||||
|
74d1dac84ecc: Download complete
|
||||||
|
99ba982a9142: Pull complete
|
||||||
|
d6b1b89eccac: Pull complete
|
||||||
|
2780920e5dbf: Pull complete
|
||||||
|
7c12895b777b: Pull complete
|
||||||
|
3214acf345c0: Pull complete
|
||||||
|
52630fc75a18: Pull complete
|
||||||
|
dd64bf2dd177: Pull complete
|
||||||
|
b839dfae01f6: Pull complete
|
||||||
|
ebddc55facdc: Pull complete
|
||||||
|
bdfd7f7e5bf6: Pull complete
|
||||||
|
2d4d7adf6272: Pull complete
|
||||||
|
40008157d8d2: Pull complete
|
||||||
|
bd8962e29291: Pull complete
|
||||||
|
cac2ae0193cb: Pull complete
|
||||||
|
74d1dac84ecc: Pull complete
|
||||||
|
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
|
||||||
|
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
|
||||||
|
docker.io/cloudflare/cloudflared:2026.6.0
|
||||||
|
1.3.3-stable: Pulling from gtstef/filebrowser
|
||||||
|
6a0ac1617861: Pulling fs layer
|
||||||
|
ef8806083e82: Pulling fs layer
|
||||||
|
b74107c861c7: Pulling fs layer
|
||||||
|
adc935def003: Pulling fs layer
|
||||||
|
4f4fb700ef54: Pulling fs layer
|
||||||
|
18695ccc900a: Pulling fs layer
|
||||||
|
45d119d5c397: Pulling fs layer
|
||||||
|
dac52db4fc51: Pulling fs layer
|
||||||
|
6d598f86b2f2: Pulling fs layer
|
||||||
|
8aa349c8396c: Pulling fs layer
|
||||||
|
45d119d5c397: Waiting
|
||||||
|
dac52db4fc51: Waiting
|
||||||
|
6d598f86b2f2: Waiting
|
||||||
|
8aa349c8396c: Waiting
|
||||||
|
adc935def003: Waiting
|
||||||
|
4f4fb700ef54: Waiting
|
||||||
|
18695ccc900a: Waiting
|
||||||
|
6a0ac1617861: Verifying Checksum
|
||||||
|
6a0ac1617861: Download complete
|
||||||
|
b74107c861c7: Verifying Checksum
|
||||||
|
b74107c861c7: Download complete
|
||||||
|
adc935def003: Verifying Checksum
|
||||||
|
adc935def003: Download complete
|
||||||
|
4f4fb700ef54: Verifying Checksum
|
||||||
|
4f4fb700ef54: Download complete
|
||||||
|
45d119d5c397: Verifying Checksum
|
||||||
|
45d119d5c397: Download complete
|
||||||
|
18695ccc900a: Verifying Checksum
|
||||||
|
18695ccc900a: Download complete
|
||||||
|
6a0ac1617861: Pull complete
|
||||||
|
dac52db4fc51: Verifying Checksum
|
||||||
|
dac52db4fc51: Download complete
|
||||||
|
8aa349c8396c: Verifying Checksum
|
||||||
|
8aa349c8396c: Download complete
|
||||||
|
6d598f86b2f2: Verifying Checksum
|
||||||
|
6d598f86b2f2: Download complete
|
||||||
|
ef8806083e82: Verifying Checksum
|
||||||
|
ef8806083e82: Download complete
|
||||||
|
ef8806083e82: Pull complete
|
||||||
|
b74107c861c7: Pull complete
|
||||||
|
adc935def003: Pull complete
|
||||||
|
4f4fb700ef54: Pull complete
|
||||||
|
18695ccc900a: Pull complete
|
||||||
|
45d119d5c397: Pull complete
|
||||||
|
dac52db4fc51: Pull complete
|
||||||
|
6d598f86b2f2: Pull complete
|
||||||
|
8aa349c8396c: Pull complete
|
||||||
|
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
|
||||||
|
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
|
||||||
|
docker.io/gtstef/filebrowser:1.3.3-stable
|
||||||
|
1.1.0: Pulling from admin/felhom-samba
|
||||||
|
897d797d2723: Pulling fs layer
|
||||||
|
3051591aa250: Pulling fs layer
|
||||||
|
ce57a3f93416: Pulling fs layer
|
||||||
|
fb94eeec2fe1: Pulling fs layer
|
||||||
|
fb94eeec2fe1: Waiting
|
||||||
|
ce57a3f93416: Verifying Checksum
|
||||||
|
ce57a3f93416: Download complete
|
||||||
|
fb94eeec2fe1: Verifying Checksum
|
||||||
|
fb94eeec2fe1: Download complete
|
||||||
|
897d797d2723: Verifying Checksum
|
||||||
|
897d797d2723: Download complete
|
||||||
|
3051591aa250: Verifying Checksum
|
||||||
|
3051591aa250: Download complete
|
||||||
|
897d797d2723: Pull complete
|
||||||
|
3051591aa250: Pull complete
|
||||||
|
ce57a3f93416: Pull complete
|
||||||
|
fb94eeec2fe1: Pull complete
|
||||||
|
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
|
||||||
|
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||||
|
gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||||
|
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
|
||||||
|
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
|
||||||
|
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
|
||||||
|
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
|
||||||
|
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
|
||||||
|
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
|
||||||
|
[golden] identity-clean + minimize …
|
||||||
|
[golden] stop + archive …
|
||||||
|
INFO: including mount point rootfs ('/') in backup
|
||||||
|
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||||
|
INFO: archive file size: 626MB
|
||||||
|
INFO: Finished Backup of VM 9100 (00:00:30)
|
||||||
|
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_08_22-18_37_25.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
|
||||||
|
[golden] publishing golden (657271965 bytes, sha256 cb439418c7005ce0…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.220.2/golden.tar.zst
|
||||||
|
[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||||
|
[golden] upload OK (HTTP 201)
|
||||||
|
GOLDEN_VERSION=0.220.2
|
||||||
|
GOLDEN_SHA256=cb439418c7005ce01bcb8126bb6385688c2f408c3c4649f6740001bc2c864ed5
|
||||||
|
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.220.2 / cb439418c7005ce01bcb8126bb6385688c2f408c3c4649f6740001bc2c864ed5
|
||||||
|
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
|
||||||
Reference in New Issue
Block a user