Files
felhom-agent/REPORT.md
T
2026-06-10 15:02:16 +02:00

54 lines
3.0 KiB
Markdown

# REPORT — slice 8B.2 (agent half): quiesce downtime optimization — `snapshotted` phase (v0.13.0) (2026-06-10)
> Overwrite-latest report. Cumulative history: [CHANGELOG.md](CHANGELOG.md). Implements the agent
> half of `TASK — Slice 8B.2`. The controller early-resume is `felhom-controller` v0.38.0. No hub change.
## Outcome
In snapshot mode an LXC vzdump only needs the app-stopped state captured at the **storage-snapshot
moment**; after that it reads from the snapshot and the app can safely resume. The agent now emits a
**`snapshotted`** phase on `GET /backup/status` when the snapshot is taken, so the controller resumes
its app early — app downtime drops from *whole-backup* to *until-snapshot* with no loss of
app-consistency.
## Phase 0 (validated first, on the demo — PVE 9.2.2)
- **Snapshot mode is in effect** on local-lvm (lvmthin) — `vzdump --mode snapshot` logs `backup mode:
snapshot`, not downgraded to stop.
- **Marker:** `INFO: create storage snapshot 'vzdump'` — the vzdump-level step after which the backup
reads from the snapshot (generic across snapshot-capable storage; only appears in snapshot mode).
- **Downtime delta:** snapshot created in the first ~1s; archive runs ~23s after. So 8B (resume at
`done`) = ~24s app-down; 8B.2 (resume at `snapshotted`) = ~1s — **~95% cut** for a 934 MB guest;
the delta grows with guest size.
## What landed
- **`BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)`** — while the vzdump runs, a watcher
tails the task log (`TaskLogTail`) for the `create storage snapshot` marker and fires `onSnapshot`
**once**. It bails on `backup mode: stop` and the marker never appears in stop mode, so it **never
fires in stop/downgraded mode**. `Backup` keeps its signature (scheduler/selftest); both share one
body. `snapshotWatchInterval` is a package var (prod 1s; tests shrink it).
- **`/backup/status` phase `snapshotted`** (between `running` and `done`): `handleBackup` passes the
hook → `markSnapshotted` flips the running job. `done`/`failed` unchanged; `snapshotted` is additive.
## Tests
`go test ./...` green; `-race` green (build server). localapi: snapshot mode → `snapshotted` before
`done`; stop mode → `snapshotted` never emitted. runner: the watcher fires on the marker; stop-mode
log never fires.
## Live validation (demo-felhom)
End-to-end on a provisioned guest (controller v0.38.0 + a postgres stack): the agent logged `backup
reached snapshotted (app may resume)` mid-backup, the controller resumed the app at that point, and
the backup proceeded to `done`. **App downtime ~3s** (quiesce→snapshotted) vs **~23s** if it had
waited for `done` (~87% cut). The snapshot backup restored **clean** (`database system was shut down`,
no WAL replay) — the early resume did not compromise app-consistency. Deployed to the demo service as
v0.13.0.
## Deferred / dependency
Snapshot-capable storage (lvm-thin/ZFS) is required for the win; on stop/downgraded storage the
controller falls back to resume-at-`done` (8B). No consistency-contract or crash-safety change. No
secrets committed.