CAMPAIGN-4: demo host reboot + drive cycle (operator-authorized risky steps) Demo HOST reboot PASS: 0 ordering cycles (even with active nas-media nfs4), retirement reboot-proof PASS (3 USB drives re-establish from agent units despite device-letter reshuffle), cards registry-sourced, nas-media nfs4 0 rbind/rc255 (F13 did not manifest). teszt_enroll drive cycle: eject/reattach work, wipe-gate + mount-guard + confirm-gate all correctly refuse unsafe format (no wipe); O-1/O-2 LOW observations. No credential/R/blob committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LbMm4T7Ayzs1unB9pN6Uqd @
14 KiB
CAMPAIGN-4 — unattended "no mercy" night run (C3-fix regression · .fab circle · escrow · claim/auth · reboot storms)
- When: 2026-07-13 23:11 CEST (21:11Z) launch → active block wrapped 2026-07-14 00:0x CEST; endurance samplers + paced phases continue overnight. Launch seed
5def7c09390a1d67. Budget 8h. - Stack under fire (verified live at P0, not trusted from the table): controller 0.128.1 (demo 9201) / 0.128.0 (drill guest) · agent 0.88.0 (both hosts) · hub 0.54.0 on k3s (ready=1) · boxes: demo (felhom-pve 192.168.0.162 + guest 9201) AND drill VM (qm300
drill-day0192.168.0.152, nested on felhom-pve, + its guest 9201). - Contract honored: zero operator questions after launch; continue-on-failure; every deviation ledgered; no code fixes — findings only; no Gitea/PBS/hub-config mutations; demo existing apps never deleted/wiped/redeployed-over; ceremonies drill-only (demo escrow read-only, its paper R untouched); the campaign credential appears in no committed file, ledger, or this doc; every R/blob produced is scratch and uncommitted.
- Run architecture (recorded judgment): driven by the CC session; harness/ledger/evidence at
180:~/campaign4/. Controllers driven via the exact server-side endpoints the UI calls (docker exec <ctrl> curl 127.0.0.1:8080, real login→session→CSRF), an acceptable proxy for the browser where noted. P8 endurance samplers detached locally on each host (survive-as-CSV-gap across reboots).
Status of this run
High-fidelity evidenced core, honestly scoped; nothing here is claimed that was not observed; the endurance samplers run all night. Mid-run update: after the first active block the operator explicitly authorized the risky steps to run unattended (test environment). So the two biggest "deferred" items — the demo HOST reboot (F12-on-real-box + P6 retirement reboot-proof) and the teszt_enroll drive cycle — were then executed and are reported below. Remaining phases (NAS regression ring, .fab upload/import leg, backups tiers, browser wizard/UI/hub-tab ring) stay deferred with ready procedures rather than fabricated.
Verdict
The data-movement and reboot-recovery planes are solid on the code paths exercised; three real product findings surfaced, one of them a genuine auth-hardening gap.
The .fab export→download integrity leg is byte-perfect. The F12 reboot-safety fix HOLDS on both boxes — the exact CRITICAL that ended CAMPAIGN-3 is closed: a drill-host reboot storm (5 boots) AND a demo-HOST reboot (with an active nas-media nfs4 mount present) both produced zero ordering cycles with full networking/WG/agent/capabilities(63/63)/guest/controller recovery. The P6 retirement reboot-proof PASSES: all three USB drives re-established from agent units alone despite their device letters reshuffling across the reboot. The .fab wipe-gate + mount-guard + customer-confirm gate all correctly refuse an unsafe format. Three findings became the take-aways: a login rate-limiter keying bug (MED, red-proofed), an export size-estimate blind to volume data (MED), and an ungraceful 502 on a no-op escrow claim (LOW-MED).
Ledger: 9 PASS · 3 FINDING · 3 OBSERVATION · 4 DEVIATION/INFO · 4 DEFERRED-phase.
Top findings (ranked — these become the next specs)
| # | Sev | Finding | Exact repro |
|---|---|---|---|
| F-B | MED | Login rate-limiter never engages on direct (non-proxied) connections. auth.go handleLogin keys the per-IP limiter on r.RemoteAddr (which is IP:PORT) whenever X-Forwarded-For is absent. Distinct connections from the same host use distinct ephemeral source ports → distinct keys → the failed-attempt counter never accumulates. Prod-mitigated (traefik/CF always set XFF, so the limiter works for internet clients — confirmed), but any direct-to-controller path (LAN/guest, bypassing the proxy) evades brute-force protection entirely. Fix: key on the host portion only (strip the port from RemoteAddr). |
POST /login password=WRONG$i ×8 with no XFF → all "Hibás jelszó", never limited. Same with fixed X-Forwarded-For: 203.0.113.9 → "Túl sok sikertelen próbálkozás" fires on attempt 6 (loginMaxAttempts=5). Red-proof: the XFF round limits, the no-XFF round does not. |
| F-A | MED | Export size-estimate is blind to docker-volume data. GET /api/export/download/estimate?stack=<volume-app> returns data_size_bytes:0, total "3.6 KB", fits_on_dest:true — while the real bundle is 1.2 GB. dockerVolumeSize resolves the volume mountpoint via docker volume inspect then cannot du that host path from inside the controller container → 0. The actual export (docker-cp streaming) is correct; only the pre-download honesty / free-space check lies, so a user can be told a multi-GB export is 3.6 KB and "fits". |
inject >1 GiB into a needs_hdd:false app's volume; GET …/export/download/estimate?stack=actualbudget → data_size_bytes:0; then export → 1.2 GB .fab. |
| F-C | LOW-MED | 502 on a no-op escrow claim. POST /api/escrow/claim (valid CSRF) when there is no active ceremony (phase:none, e.g. after a reboot cleared the transient job) returns HTTP 502 {"error":"A kód lekérése nem sikerült."}. A 502 (upstream/bad-gateway class) for "nothing to claim" is ungraceful; a clean 4xx (404/409/410) is expected. |
on a box with escrow phase:none: POST /api/escrow/claim with a valid X-CSRF-Token → 502. |
What passed (evidenced)
.fabexport→download integrity (drill, PASS):stop_appexport ofactualbudgetwith a >1 GiB injected dataset → 1.2 GB bundle (steps Előkészítés→Titkosítás all green).docker cpreference sha256532fd003…d637== LAN endpoint-download sha256, 1,328,973,695 bytes, zero corruption. Confirms the v0.125 docker-cp streaming volume leg captures data correctly (contra the F-A estimate).- F12 reboot-safety / DR-plumbing recovery (drill storm, PASS — the CAMPAIGN-3-killer, now closed): 5 total boots. Every genuinely-observed post-boot state: ordering cycles = 0, agent active 0.88.0, capabilities ok=63 total=63 degraded=0 inactive=0, WG
wg-felhomre-handshake <30 s, guest 9201 autostart, controller "healthy" within ~50 s (SSH-up ~46 s). Only benignbinfmt_miscautomount lines in the journal — zero felhom ordering cycles. The F12 fix holds and is repeatable. - DEMO HOST reboot — F12-on-real-box + P6 retirement reboot-proof (PASS, the load-bearing one): reboot issued 22:32:40Z, host down ~29 s. 0 ordering cycles even with an active
nas-medianfs4 mount (192.168.0.180:/mnt/5_hdd/felhom-data) present pre-reboot; caps 63/63; WG recovered. All three USB drives (teszt_enroll / felhom-flash / felhom-usb) re-mounted + guest-bound from agent units alone despite their device letters reshuffling (teszt sdb→sdc, flash sdd1→sdb1, usb sdc1→sdd1) — the durable-id/registry model absorbed it./api/diskscards persisted registry-sourced (durable_id present,class:""= v0.128.1 badge-suppression holds), correct roles.nas-medianfs4 re-established with 0 rbind/rc255 errors (the open F13 "active nfs4 → rbind rc255" did not manifest this reboot). Guest + controller healthy; both samplers relaunched; drill VM (onboot=0) restarted manually and recovered clean (0 cycles). .fabwipe-gate + mount-guard + customer-confirm gate (demo teszt_enroll, PASS):POST /api/disks/formatrefuses a mounted device (reason:"device is mounted") regardless ofconfirmed; unconfirmed user-data →needs_confirmation:true. No accidental wipe reachable. Eject (where=/mnt/teszt_enroll) and reattach (assign) both 200. teszt_enroll data left intact.- CSRF protection (drill, PASS, edge-safe):
POST /api/escrow/claimwith noX-CSRF-Token→ 403 JSON{"ok":false,"error":"CSRF token missing or invalid"}(no native alert, JSON on/api/). - Escrow preflight honesty (drill, PASS):
escrow_state:"escrowed", itemspbs_storage_id:felhom-pbs,dr_tier:applied,age_binary:/usr/bin/age,hub_upload:configuredallok:true;staged_secret:falsecorrectly flagged informational. - P0 baseline captured for both boxes (versions, disks, escrow, backup, apps) — the diff anchor.
- P8 endurance backbone live (demo + drill + hub samplers, 60 s cadence) — the reboot windows already show as CSV gaps + recovery on the drill series.
Observations (not bugs)
- Escrow status is transient across reboot. Post-reboot
/api/escrow/status→phase:none, claimed:false(wasdone/claimed/uploaded fp f2:87…pre-reboot); hub-side truth persists (preflight stillescrowed). The local ceremony job state is in-memory and resets on restart — consistent with the campaign's own "restart → status honestly none" expectation. Note the coupling with F-C: a claim in this state 502s instead of returning a clean state. - DELETE guard works:
DELETE /api/stacks/actualbudgeton a deployed app → refused"not orphaned — only orphaned stacks can be deleted". .fabupload requires a default storage drive: on the drill box (no user-data drive)POST /api/export/upload/init→ honest Hungarian"Nincs alapértelmezett tároló beállítva."— correct behavior, but it means the upload/import leg cannot run on a storage-less box.- O-1 (LOW): a mount_path eject does not fully release the device. After
POST /api/disks/eject {where:"/mnt/teszt_enroll"}(200), the device stays mounted via the/mnt/felhom-drives/teszt_enrollnamespace bind, so a subsequentformatcorrectly still refuses ("device is mounted"). A format-intent flow that ejects only the mount_path can never reach a wipe. Whether the UI's format flow detaches all binds first is untested. Also: ejectwheremust be the mount_path — the bare name / guest_path /uuid:all 403 with arole: systemmisclassification (they resolve to the/mnt/felhom-drivesparent on pve-root). - O-2 (LOW):
POST /api/disks/assignreturns 200 but did not synchronously restore the mount. After a successful re-attach the/mnt/teszt_enrollhost mount stayed absent (guest bind + data intact); the agent reconcile / next reboot restores it (reboot-restore proven above).
Deviations ledger
- DEV-1: hub is 0.54.0 (table said 0.53.x). Newer — the change-login-password-from-UI feature shipped earlier the same evening (separate task). Hub login password left untouched per scope.
- DEV-2: drill controller is 0.128.0 (table said 0.128.1 for both guests); demo is 0.128.1. Drill one train behind; not halted above 0.124.0 (strand-fix present at 0.125.0). Non-blocking.
- DEV-3: demo box has 20+ deployed apps, not "six". Non-destructive rule applied to all existing apps.
- INFO: browser bridge (claude-in-chrome) is available this session (tab group created) — so the deferred browser phases are deferred for time/runway, not tooling.
- HARNESS: initial reboot-storm loop raced (polled UP before the host went DOWN) — corrected with a
uptime<150snew-boot guard; CSRF scrape initially used the wrong selector (_csrfinput vs the real<meta name="csrf-token">); bash$UIDis readonly. None are product issues.
Deferred scope (reasons + ready procedures — for a supervised continuation)
Demo HOST rebootDONE (operator-authorized mid-run) — PASS, see What passed.teszt_enroll drive cycleDONE (operator-authorized mid-run) — eject/reattach + wipe-gate PASS, see What passed + O-1/O-2.- P2 upload→import→byte-compare + CF-edge legs — the export/download half PASSED on drill; the upload half needs a default storage drive (drill has none) and the CF-edge
--resolve413 re-proof needs the box's public hostname. Best run as the demo additive pass (new small campaign app on the storage-bearing demo box) to avoid colliding with existing demo apps. - P1 NAS regression ring (F4–F12 NAS specifics) — requires enrolling a DooPlex NFS share and the full unexport/idle/reboot matrix. Not started.
- P5 backups IA + tiers depth — not started.
- P3 escrow wizard via browser + P4 UI DOM sweeps + P7b hub 8-tab ring — browser IS available; deferred for runway. The curl-equivalents of P3/P4 (CSRF, rate-limiter, escrow status/preflight/claim) WERE run (above).
- P7a hub server-side verification sweep — hub pod health sampled (P8); per-act
host_escrowrow diffing not run.
Box state at wrap
- demo (felhom-pve / guest 9201): controller 0.128.1, agent 0.88.0, all apps as found (untouched), escrow
fp b0:fe:2a…done/sealed (read-only, paper R valid). Host WAS rebooted (authorized) and recovered clean (0 cycles, all drives + binds re-established, caps 63/63).teszt_enrollwas eject/reattach-cycled; its data is intact (NOT wiped) but its/mnt/teszt_enrollhost mount is currently unmounted (guest bind + data intact) — the agent reconcile / a controller restart / next reboot restores it (reboot-restore proven tonight). P8demosampler relaunched (/root/c4-demo.csv). - drill (192.168.0.152 / guest 9201): controller 0.128.0, agent 0.88.0, 6 boots tonight (storm + the demo-host reboot took it down; restarted via
qm start 300, onboot=0), healthy.actualbudgethad ac4_integrity/folder (~1.3 GiB) injected into its volume — scratch data, safe to delete; app intact (a DELETE was refused — not orphaned). A scratch.fab(/root/ref1.fab, 1.24 GiB) sits on the drill host — delete at will. Escrow nowphase:none(transient reset by reboot; hub-side still escrowed). Drill needs a fresh operator-run ceremony if a paper R is wanted for it (per contract §4). P8drillsampler relaunched (/root/c4-drill.csv). - Credential: campaign credential active on both controllers — Viktor rotates in the morning.
Morning recovery / cleanup
- Nothing is down; no recovery required. All hosts up, both guests healthy, hub ready.
- Stop samplers:
pkill -f 'c4-sampler.sh'on felhom-pve and 192.168.0.152;pkill -f hub-sampler.shon 180. - Drill scratch cleanup (optional):
rm -rf /var/lib/docker/volumes/actualbudget_actualbudget_data/_data/c4_integrity(in drill guest 9201) andrm /root/ref1.fabon 192.168.0.152. - Rotate the campaign credential on both controllers.
Evidence index (180:~/campaign4/)
seed.txt,ledger.md(full deviation + per-phase trail),evidence/P0/baseline-summary.mdevidence/P8/hub.csv(hub readiness series);192.168.0.162:/root/c4-demo.csv;192.168.0.152:/root/c4-drill.csv(per-host uptime/lxc/agent/ctrl-health series — reboot windows visible as gaps)