Files
felhom-controller/REPORT.md
T
admin c6b69d888e
gates / gates (push) Successful in 21s
v0.205.0 — a run that skipped an app the customer selected is not successful (R-234)
THE VERDICT. The R-203 block already said "a warning beside a success is read as a
success" and applied it to ONE of the two shapes it describes: an app missing a
declared mandatory FOLDER made the run incomplete, while an app skipped ENTIRELY
still reported ok. Both do now. Which skips count, decided by measurement:
selected+deployed with no recovery unit YES; selected but NOT deployed no (named,
with what to do — a box left amber by an app somebody removed is a status nobody
reads); disconnected/decommissioned drive no (own signal); nothing selected no.
LastSuccess and SnapshotCount still record what WAS captured.

THE FILED MECHANISM WAS NOT THE MEASURED CAUSE, and saying so is the point. §3
stated that toggling an app on leaves it without a bundle so the first run skips
it. Measured on demo-hp: the run's own pre-dump phase calls captureAllRecoveryUnits
for every DEPLOYED stack, through admitApp, before the push — a unit moved aside
was RECREATED and the run reported ok. That state does not survive a run.

What actually produced the 2026-08-06 sequence: the manual run was dropped by the
single-flight while an earlier run was still going. runOffboxBackup returned nil,
the handler had already answered "A tavoli mentes elindult", and the card then
showed the PREVIOUS run's green verdict — read as covering the app just selected.
The decision is now taken synchronously in the handler and a dropped request says
so. The nightly path still returns nil on purpose: nobody asked, and it retries.

§7.3 measured before deciding: CaptureRecoveryUnit writes a few KB of compose +
manifest, only ENUMERATES dumps rather than creating them, is idempotent and does
NOT stop the app — and already runs inside the off-site run. So there is no wait to
remove for a deployed app and NOTHING was built.

28 packages ok, 9/9 gates. Four red-proofs, each asserted to have applied. Fixture
note: the shared provider's ListDeployedStacks returned nil, so Scenario A first
passed for the wrong reason; fixed with an opt-in deployed set that defaults to nil.
2026-08-06 21:58:21 +02:00

7.8 KiB

REPORT — v0.205.0: a backup that skipped an app the customer chose is not „Rendben" (R-234)

2026-08-06. Controller v0.204.0 → v0.205.0. MinAgent unchanged (0.127.0). felhom-agent, app-catalog-felhom.eu untouched. No hub change.

The correction that outranks the task

§3 stated the mechanism as established: "the off-site copy sends the local recovery bundle an app already has; switching an app on does not create one, so the first run after the toggle finds nothing to send." Measured on demo-hp, that is not what the code does.

The off-site run's own pre-dump phase calls captureAllRecoveryUnits(), which writes a unit for every deployed stack — through admitAppbefore the per-app push loop. I moved calibre-web's unit aside (never deleted) on demo-hp and triggered a run through the real endpoint:

BEFORE : units = calibre-web, opengist, privatebin   status=ok  snapshots=9
moved  : calibre-web unit → .calibre-web.R234-aside
RUN    : POST /backup/offbox/run → 302, finished ~90 s, status=ok
AFTER  : status=ok  snapshots=9  warning=(none)
unit   : /mnt/sys_drive/felhom-data/backups/primary/calibre-web  → RECREATED

So "selected but no bundle yet" does not survive a run for a deployed app. The filed mechanism could not have produced the 2026-08-06 sequence.

What did. POST /backup/offbox/run launched the run in a goroutine and immediately answered „A távoli mentés elindult". Inside, acquireRunning refused (a run was already in flight) and runOffboxBackup returned nil — no error, no signal. The card then showed the previous run's „✓ Rendben · 1 pillanatkép", read as covering the app just selected. It did not: the restore refused for that app minutes later, and a third run carried it. That is R-234's real cause, and it is the same family — a skip that reached a log and not a verdict.

Live validation (§12)

# what observable
1 selected app with no unit → run the run recreated the unit and reported ok — the measurement above. The verdict change is exercised by the run-level tests, because production cannot easily be held in that state
2 counters intact snapshot_count 9 → 9, last_success advanced to 2026-08-06T19:34:27Z — what was captured is still recorded
3 the third-run behaviour arriving sooner not applicable as filed: there is no wait for a deployed app; the unit is created by the run itself (§7.3)
4 a box with nothing selected unchanged — Scenario D is pinned by test, and demo-hp's other apps stayed ok

Access note: demo-hp's dashboard password had been changed by the customer claim, so I reset it through the documented --print-reset-code escape hatch (code streamed file→file, extracted by shape, shredded; the new password is a 24-char generated value in a 0600 file on DooPlex). That is a deliberate change to a Tier-0 demo box, recorded here.

§7.2 — which skips count, as decided

skip counts? why
selected + deployed, no recovery unit YESincomplete the app the customer chose is not in the snapshot at all
selected but not deployed no — named, with what to do a box left amber forever by an app somebody removed is a status nobody reads
drive disconnected / decommissioned no it has its own card and its own signal; re-reporting it here would double-count
nothing selected at all no unchanged: the existing „nincs mentésre jelölt alkalmazás" notice

§7.3 — the measurement, and the decision it forced

CaptureRecoveryUnit writes compose/ (docker-compose.yml, .felhom.yml, app.yaml) + manifest.jsona few KB, per admission.go's own note — and enumerates the DB/volume dumps already present rather than creating them. It is idempotent (skips all writes when current) and does not stop the app; the stopping work belongs to the separate dump flow.

Decision: nothing to build. The cheap, non-stopping capture already runs for every deployed stack inside the off-site run's pre-dump phase, through admitApp. For a deployed app there is no wait to remove, and the §7.3 branch that would have added an inline capture would have duplicated it.

§7.4 — every changed Hungarian string

  • „Ezek az alkalmazások NEM kerültek be a távoli mentésbe, mert még nincs helyi mentési egységük: %s. A következő mentés általában már elkészíti — ha a második futás után is itt szerepelnek, szólj az üzemeltetőnek."
  • „Ezek az alkalmazások ki vannak jelölve távoli mentésre, de nincsenek telepítve, ezért nem menthetők: %s. Ha már nincs rájuk szükséged, vedd ki a kijelölésüket a Távoli mentés oldalon."
  • „Már fut egy távoli mentés — ez a kérés nem indított újat. A most látható eredmény még a korábbi futásé; várd meg, míg ez befejeződik."
  • sibling, extended so both read alike: „Figyelmeztetés: a(z) %s alkalmazás egyes adatmappái nem kerültek a távoli mentésbe: %s. Ellenőrizd, hogy a mappák megvannak-e a meghajtón; ha igen és ez a következő mentés után is látszik, szólj az üzemeltetőnek."

The replaced sentence was: „Figyelmeztetés: %d alkalmazásnak nincs elérhető mentése, ezek kimaradtak: %s" — a count with no next step, reading the same whether the customer must act or simply wait.

Tests and red-proofs

go build · go vet · go test ./...28 packages ok. controller_gates.py --fast9/9 OK.

New: internal/backup/offbox_verdict_r234_test.go (Scenarios A, C, D, F — run-level) and internal/web/offbox_run_inflight_test.go (handler-level).

# mutation (each asserted to have applied) result
RP-1 drop unprotected from the verdict — the defect itself SkippedSelectedAppIsIncomplete FAIL — „LastStatus = ok, want incomplete"
RP-2 remove the classification switch; count every skip SelectedButUndeployedIsNamedNotCounted FAIL — a removed app turns the box amber
RP-3 stop adding skipped apps to the operator signal SkippedSelectedAppIsIncomplete FAIL — „operator signal … got map[]"
RP-4 remove the synchronous in-flight check in the handler InFlightRequestIsNotReportedAsStarted FAIL

All restored; grep -c RED-PROOF = 0 in both files afterwards; suite green again.

Fixture note: the shared offbox3aProvider.ListDeployedStacks() returned nil, so my first run of Scenario A passed for the wrong reason and F passed vacuously. Fixed with an opt-in deployed map that defaults to nil, so no existing fixture's behaviour moves.

Files

  • internal/backup/offbox.goErrOffboxRunInFlight, offboxWholeUnitGap, the skip classification, missingUnprotected/missingNotDeployed, stackDeployed, driveUnavailableFor, the verdict, the messages
  • internal/backup/offbox_capture.go — the sibling message
  • internal/backup/backup.goAcquireRunningForTest/ReleaseRunningForTest (test-only seams)
  • internal/web/offbox_handlers.go — the synchronous single-flight refusal
  • new internal/backup/offbox_verdict_r234_test.go, internal/web/offbox_run_inflight_test.go
  • internal/backup/offbox_3a_test.go — opt-in deployed set
  • CHANGELOG.md, controller/README.md, CONTEXT.md

Observations, not acted on

  • The zero-selection notice reads „Sikeres — nincs mentésre jelölt alkalmazás" while the status is ok. It is honest, but "Sikeres" beside "nothing is covered" is the same rhetorical shape this session is about, one notch weaker. Not touched: Scenario D forbids changing that path.
  • res.missing is still used verbatim for the no-silent-success error text (every app missing → error). Correct as-is, and deliberately left, since that path already refuses loudly.