records-carried: controller v0.276.0 (R-697 closed, R-700 filed), 07 §6.6, register 336 -> 336, STATUS
gates / gates (push) Successful in 26s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-27 18:04:00 +02:00
parent b52cb6ec94
commit 06744dbacc
14 changed files with 128 additions and 14 deletions
@@ -0,0 +1,3 @@
# 9202 before: gitea.dooplex.hu/admin/felhom-controller:0.275.0
# 9202 after: gitea.dooplex.hu/admin/felhom-controller:0.276.0 Up 25 seconds (healthy)
2026/09/27 16:00:45 main.go:337: [INFO] felhom-controller 0.276.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
@@ -0,0 +1,8 @@
# global floor 0.276.0 (declared MinAgent 0.131.0) — 2026-09-27T16:01:26Z
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
2026-09-27T16:01:36Z
demo-hp: gitea.dooplex.hu/admin/felhom-controller:0.276.0
demo-felhom: gitea.dooplex.hu/admin/felhom-controller:0.276.0
gitea.dooplex.hu/admin/felhom-controller:0.276.0 Up 22 seconds (healthy)
gitea.dooplex.hu/admin/felhom-controller:0.276.0 Up 25 seconds (healthy)
@@ -0,0 +1,12 @@
state running
pinned_images {"paperless-postgres": "postgres:18-alpine", "paperless-redis": "redis:7-alpine", "paperless-webserver": "ghcr.io/paperless-ngx/paperless-ngx:2.20.15"}
desired_state "running"
conversion_copy {"volume": "paperless-ngx_paperless_postgres_data", "copy": "paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z", "at": "2026-09-27T09:45:15Z", "from": 16, "to": 18, "service": "paperless-postgres"}
earlier_conversion_copies null
paperless-ngx_paperless_data
paperless-ngx_paperless_postgres_data
paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z
paperless-ngx_paperless_redis_data
2026/09/27 16:00:45 scheduler.go:102: [INFO] [scheduler] Registered periodic job: conversion-copy-release (every 1h0m0s)
2026/09/27 16:00:45 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="conversion-copy-release" interval=1h0m0s totalJobs=10
@@ -0,0 +1,46 @@
# records-carried — 2026-09-27 (controller v0.276.0): a restore and a drive move keep the app's records
Architecture read before any claim: `07-backup-architecture.md` §6.6 (now with "What a restore keeps"), `09` §6.4 part 10
(the conversion copy). Rows: R-697 (closed), R-700 (new; fixed, live proof open), R-691 (2) (not built — why in the row).
## Found
- **R-697 (known):** `PersistUnitRedeployConfig` built a fresh `AppConfig`, so a restore dropped `conversion_copy` — the
kept 16 datadir volume was never released. Also dropped: `desired_state` (a dead app after a restore then reads as
"unknown intent" and is not alarmed), `failed_update_step`, `last_update_undone`, `last_auto_update`.
- **Second orphan (same row):** after a restore to the OLD major the ladder converts again; `recordConversionCopy` overwrote
the first copy's record with the second's.
- **R-700 (new, by reading, not seen on a box):** `doFlipRedeploy` (drive move) persisted through the same fresh write, so it
ALSO dropped `pinned_images`. `sync.renderSource`: deployed + unpinned → the catalog copied verbatim → the next `up` runs
the catalog's newest version. Pin adoption runs only at controller start.
## Fixed (v0.276.0, controller `820e8ef`)
`carryLifeRecords` in the restore write; `earlier_conversion_copies` + a release loop over every kept copy;
`persistDriveFlip` (load-then-save, `HDD_PATH` only) + `upFromAppConfig`.
## Red-proofs (`redproofs/`, each seen failing, tree restored, suite green after)
| | pre-fix shape | failed at |
|---|---|---|
| RP1 | no carry in the restore write | `conversion_copy = <nil>` |
| RP2 | a newer record overwrites the older | `earlier []` |
| RP3 | the drive move persists through the restore write | `pinned_images = map[]` |
| RP4 | `doFlipRedeploy` back on `RedeployFromEnv` | the wiring test |
## Live (endpoint level + guest reads)
- `R/R1-9202-0.276.0.txt` — 9202 by hand, healthy. `R/R2-floor.txt` — global floor 0.276.0 (MinAgent 0.131.0), both demo
boxes on 0.276.0 within 10 s, healthy.
- `R/R3-9202-paperless-record.txt` — after the upgrade paperless-ngx keeps `conversion_copy` (16 → 18, at 09:45:15Z) and
its `.pre-update-20260927T094421Z` volume; the release job registered.
- **Night (N/):** the release on a real 18 dump — see `N/README.md` once written.
## Not proven live
- **R-700:** no Tier-0 guest has two drives (9202 has one: `scratch_hdd`). Unit-proven only; row stays WATCHING.
- **R-697's carry across a real restore:** restoring paperless now would bring it back at 16 (its unit's data is the
pre-conversion dump) and would spoil the night's release proof. Unit-proven only.
Teardown: provisioned nothing. Machine: 9202 controller image changed 0.275.0 → 0.276.0 (kept). Host: nothing. Hub: global
floor 0.275.0 → 0.276.0.
@@ -0,0 +1,11 @@
RP1: carryLifeRecords call removed from PersistUnitRedeployConfig (= v0.275.0 restore write)
controller/internal/stacks/deploy.go | 11 ++++
controller/internal/stacks/migrate.go | 47 ++++++++++----
controller/internal/stacks/pgconvert.go | 109 +++++++++++++++++++++++---------
3 files changed, 126 insertions(+), 41 deletions(-)
2026/09/27 17:49:58 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
--- FAIL: TestR697_ARestoreKeepsTheConversionCopyRecordSoTheCopyIsReleased (0.01s)
r700_records_carried_test.go:43: after the restore conversion_copy = <nil>, want nextcloud_pgdata.pre-update-20260913T100000Z kept
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s
FAIL
@@ -0,0 +1,6 @@
RP2: recordConversionCopy overwrites the older record (= v0.275.0)
--- FAIL: TestR697_ASecondConversionDoesNotOrphanTheFirstCopy (0.00s)
r700_records_carried_test.go:74: records: current &{Volume:nextcloud_pgdata Copy:nextcloud_pgdata.pre-update-B At:2026-09-13T11:00:00Z From:16 To:18 Service:db}, earlier [] — want B current and A kept as earlier
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.010s
FAIL
@@ -0,0 +1,6 @@
RP3: persistDriveFlip = PersistUnitRedeployConfig(env with HDD_PATH=target) (= v0.275.0's write, plus this release's carry)
--- FAIL: TestR700_ADriveMoveKeepsThePinAndTheRecords (0.00s)
r700_records_carried_test.go:109: pinned_images = map[] — the moved app is unpinned and will jump to the catalog's version
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s
FAIL
@@ -0,0 +1,13 @@
RP4: doFlipRedeploy back on RedeployFromEnv
// the pin and every record. This used to be RedeployFromEnv, whose fresh app.yaml dropped the pin: the
// syncer then copied the catalog verbatim and the next start jumped the app past its ladder.
if err := m.RedeployFromEnv(name, nil); err != nil { // RP4
return err
}
if !m.waitHealthy(name) {
return util.MsgError("err.stacks.az_alkalmazas_nem_indult_el_az")
}
return nil
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.008s
FAIL