Files
felhom.eu/documentation/audits/evidence-spike-restic-restore-2026-08-31/28-q5-schedule.txt
T
admin 130f7a6eba
gates / gates (push) Failing after 17s
R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.

Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.

Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.

Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.

Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).

Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.

Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.

Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.

RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.

Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.

Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.

Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).

golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.

Ceiling R-406 -> R-409.
2026-08-31 15:59:09 +02:00

112 lines
8.0 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
=== scheduled off-site jobs (main.go) ===
663: sched.Every("status-refresh", 10*time.Second, func(ctx context.Context) error {
666: sched.Every("stack-scan", 2*time.Minute, func(ctx context.Context) error {
669: sched.Every("health-probes", 10*time.Second, func(ctx context.Context) error {
679: sched.Every("system-health", healthInterval, func(ctx context.Context) error {
726: sched.Every("deadapp-check", 30*time.Second, func(ctx context.Context) error {
742: sched.Every("ring-spill", 30*time.Second, func(ctx context.Context) error {
876: sched.Daily("db-dump", dbLeg, func(ctx context.Context) error {
914: sched.Every("offsite-credential-retry", 5*time.Minute, func(ctx context.Context) error {
937: sched.Every("backup-cache", 5*time.Minute, func(ctx context.Context) error {
962: sched.Daily("tier2-backup", tier2Leg, func(ctx context.Context) error {
1095: sched.Daily("offbox-backup", offboxLeg, func(ctx context.Context) error {
1109: sched.Daily("offsite-abandon-sweep", "05:10", func(ctx context.Context) error {
1139: sched.Daily("offsite-integrity", "06:00", func(ctx context.Context) error {
1147: sched.Daily("metrics-prune", "04:00", func(ctx context.Context) error {
1198: sched.Daily("fill-watch", "03:30", func(ctx context.Context) error { return fillWatcher.Check() })
1228: sched.Every("hub-report", pushInterval, func(ctx context.Context) error {
1256: sched.Every("selfupdate-check", checkInterval, func(ctx context.Context) error {
1266: sched.Daily("selfupdate-auto", cfg.SelfUpdate.AutoUpdateTime, func(ctx context.Context) error {
1293: sched.Daily("asset-sync", cfg.Assets.SyncSchedule, func(ctx context.Context) error {
1464: sched.Every("geo-verify", 6*time.Hour, func(ctx context.Context) error {
sync_schedule: "05:00"
backup:
db_dump_schedule: "02:30"
enabled: true
prune_schedule: weekly
restic_password_file: /opt/docker/felhom-controller/data/restic-password
restic_schedule: "03:00"
retention:
keep_daily: 7
keep_monthly: 6
keep_weekly: 4
customer:
domain: enkisfelhom.hu
email: doodoo21@freemail.hu
id: demo-hp
name: Demo HP
telegram_chat_id: ""
git:
branch: main
--
health_check_schedule: "06:00"
healthchecks_base: https://status.felhom.eu
system_health_interval: 5m
thresholds:
backup_max_age_hours: 36
cpu_warn_percent: 90
disk_crit_percent: 90
disk_warn_percent: 80
memory_warn_percent: 85
temperature_warn_celsius: 75
internal/backupwindow/backupwindow.go:54:func LegTimes(start string) (db, tier2, offbox string) {
internal/backupwindow/backupwindow.go-55- m, err := ParseHHMM(start)
internal/backupwindow/backupwindow.go-56- if err != nil {
internal/backupwindow/backupwindow.go-57- return "", "", ""
internal/backupwindow/backupwindow.go-58- }
internal/backupwindow/backupwindow.go-59- return FmtHHMM(m), FmtHHMM(m + tier2OffsetMin), FmtHHMM(m + offboxOffsetMin)
internal/backupwindow/backupwindow.go-60-}
internal/backupwindow/backupwindow.go-61-
internal/backupwindow/backupwindow.go-62-// GateWindow returns the whole-guest backup gate bounds [W+2h, W+6h) as HH:MM strings (for the UI
internal/backupwindow/backupwindow.go-63-// "kb. <from>–<to> között" line and the gate-denial log). Empty strings on an invalid start.
internal/backupwindow/backupwindow.go-64-func GateWindow(start string) (from, to string) {
internal/backupwindow/backupwindow.go-65- m, err := ParseHHMM(start)
internal/backupwindow/backupwindow.go-66- if err != nil {
internal/backupwindow/backupwindow.go-67- return "", ""
internal/backupwindow/backupwindow.go-68- }
internal/backupwindow/backupwindow.go-69- return FmtHHMM(m + gateStartMin), FmtHHMM(m + gateEndMin)
internal/backupwindow/backupwindow.go-70-}
internal/backupwindow/backupwindow.go-71-
internal/backupwindow/backupwindow.go-72-// EffectiveWindow resolves the active window by precedence: a valid settings value wins over a valid
internal/backupwindow/backupwindow.go-73-// controller.yaml value, which wins over DefaultWindow. An empty or corrupted value simply falls
internal/backupwindow/backupwindow.go-74-// through — so a bad settings string degrades to the yaml default rather than breaking scheduling.
internal/backupwindow/backupwindow.go-75-func EffectiveWindow(settingsVal, yamlVal string) string {
internal/backupwindow/backupwindow.go-76- if Valid(settingsVal) == nil {
internal/backupwindow/backupwindow.go-77- return settingsVal
internal/backupwindow/backupwindow.go-78- }
internal/backupwindow/backupwindow.go-79- if Valid(yamlVal) == nil {
internal/backupwindow/backupwindow.go-80- return yamlVal
internal/backupwindow/backupwindow.go-81- }
internal/backupwindow/backupwindow.go-82- return DefaultWindow
internal/backupwindow/backupwindow.go-83-}
10:// DefaultWindow is the last-resort window when neither settings nor controller.yaml supplies one.
12:const DefaultWindow = "02:30"
17: tier2OffsetMin = 60 // tier-2 mirror at W+60m
18: offboxOffsetMin = 105 // off-box copy at W+105m
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: status-refresh (every 10s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="status-refresh" interval=10s totalJobs=1
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="stack-scan" interval=2m0s totalJobs=2
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: health-probes (every 10s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="health-probes" interval=10s totalJobs=3
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: system-health (every 5m0s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="system-health" interval=5m0s totalJobs=4
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="deadapp-check" interval=30s totalJobs=5
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="ring-spill" interval=30s totalJobs=6
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-01 02:30 CEST
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="db-dump" schedule="02:30" nextRun=2026-09-01T02:30:00+02:00 totalJobs=7
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="offsite-credential-retry" interval=5m0s totalJobs=8
2026/08/31 12:33:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="backup-cache" interval=5m0s totalJobs=9
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-01 03:30 CEST
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="tier2-backup" schedule="03:30" nextRun=2026-09-01T03:30:00+02:00 totalJobs=10
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-01 04:15 CEST
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offbox-backup" schedule="04:15" nextRun=2026-09-01T04:15:00+02:00 totalJobs=11
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-01 05:10 CEST
2026/08/31 12:33:29 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-abandon-sweep" schedule="05:10" nextRun=2026-09-01T05:10:00+02:00 totalJobs=12
2026/08/31 12:33:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST
grep: write error: Broken pipe