Commit Graph

1051 Commits

Author SHA1 Message Date
admin 38d28b5b62 v0.234.0: seed installed_images at startup, so the label appears on an app nobody touched
gates / gates (push) Successful in 13s
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.

BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.

And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.

Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.

+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
2026-09-03 11:56:43 +02:00
admin 32da46cd64 REPORT: the mirror base images ARE byte-identical to Docker Hub - measured, not assumed
gates / gates (push) Successful in 13s
The throttle cleared 40 minutes later, so the check was run instead of left as a
one-command IOU. 'docker pull docker.io/library/<img>' answered 'Image is up to
date' for both bases - Docker Hub's own manifest resolved to the images already
local, the ones the mirror supplied and 0.233.0 was built from. The two manifest
indexes are also identical between registries.

Also corrected in passing: the two sha256 values recorded earlier are local IMAGE
IDs, not manifest-list digests. I conflated them once during this very check, so
the report now says which is which.
2026-09-02 21:03:24 +02:00
admin 50312b63bc REPORT: the badge is proven live too, and the stale-password claim was mine to retract
gates / gates (push) Successful in 13s
Naprakesz on /stacks (x2) and /apps/bookstack; NO badge at all on /apps/docmost,
a deployed app with no record - absent is UNKNOWN, not current; and 'Frissites
elerheto - 52 napja' on both surfaces, the age real arithmetic on bentopdf's
catalog_since. Staged by editing one compose tag with no restart and no up -d,
reverted byte-identically. ASCII fragments with positive and negative controls.

Section 7 now opens with the correction rather than burying it: I reported the
vaulted password as stale on both boxes; it was fine, and I had stripped only
double quotes from a single-quoted value. The keeper is that I read the
controller's 'Failed login' past what it discriminates.
2026-09-02 20:49:01 +02:00
admin 2fb558552a REPORT: v0.233.0 live — the record is proven on metal, the badge render is not, and why
gates / gates (push) Successful in 13s
The record: proven on demo-hp through the boot reconciler (a real production
caller, no hand-set state) on a single-service AND a multi-service app, with all
three digests matching ground truth read independently beforehand.

The badge render: NOT validated. The vaulted dashboard password is stale on BOTH
demo controllers; the five attempts are listed rather than summarised, and the
controller's own log is the discriminator that says wrong password, not wrong
host header. Filed as R-453 and raised in STATUS.md item 9.

Also recorded: the build was blocked by a Docker Hub 429 and the base images came
from Google's Hub mirror, with both digests written down so the identity check is
one command when the throttle clears - KNOWN, not measured here.
2026-09-02 20:37:01 +02:00
admin 8025304acc v0.233.0: record what each compose service actually installed, and badge whether it is current
gates / gates (push) Successful in 12s
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no
auto-update, the three lifecycle buttons byte-identical.

Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each
entry carrying ref + repo digest + first-seen timestamp. Written by
Manager.recordInstalledImages after a successful compose up from StartStack,
RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never
from docker-compose.yml: the syncer overwrites a deployed app's compose on a
15-minute cycle and the two disagreed for 25 minutes in the spike's own
measurement. A failed write NEVER refuses the action - the deliberate opposite
of SetDesiredState, because this is an observation and that is an intent. Not
called from StartStackServices (the R-47 DB-only window). Its own docker seam
with a context and a 30s timeout, which neither existing exec helper has.

Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the
recorded ref per service against what the current template pins and returns a
*MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS.
NO RECORD RENDERS NOTHING: absent means unknown and never means current. No
version number reaches the customer and no registry is queried.

Known limitation, filed not hidden: 23 catalog pins float, so those apps can read
Naprakesz when the image behind the tag has moved.

+17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real
RestartStack plus an AST walk of the four call sites. Three companion red-proofs
run and reverted.
2026-09-02 20:18:01 +02:00
admin 960d29b061 REPORT: the CI outcome - four red runs, two real causes, all now green
gates / gates (push) Successful in 12s
Adds section 11. The sharpest result of the whole sweep is in it: my own decoy-coverage gate
identified a repository by its DIRECTORY NAME and went blind the first time CI ran it, because the
act-runner checks out into a folder called hostexecutor. The gate written that morning to catch
name-for-fact was matching a name, in the first ten lines of its own main loop (R-428).

The other cause was my push ordering - two repos citing R-421 pushed before felhom.eu carried the
row - which instructions_gate convicted exactly as designed.
2026-09-01 12:48:50 +02:00
admin 22983885f1 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 13s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:54 +02:00
admin 670def61b4 REPORT: the decoy sweep - 29 gates, 16 fooled, 10 fixed, 6 honestly untested
gates / gates (push) Failing after 14s
Opens with the survey table. Records the three numbers, every live hole with its decoy and row, the
six gates no plausible decoy could be built for, the meta-gate's 20-name exemption list, and which of
the five defining rows actually closed (R-419; R-378 explicitly did NOT).

Includes my own mistakes by name - five decoys withdrawn as illegitimate, a 36-vs-44 arithmetic
artefact I announced before checking, an rc==0 read as a hole for a gate where rc is not the
question, a bash heredoc that ate my backticks, and an R-419 fix that was too strict and rejected
genuine markers until its own gate convicted this very report.
2026-09-01 12:42:30 +02:00
admin 681cc663ef decoy sweep: eight holes in this repo's gates, all measured, all fixed (R-421)
gates / gates (push) Failing after 13s
Every gate was DECOYED - the label constructed without the fact, the gate run, the verdict recorded.
No verdict here was reached by reading, because reading is exactly how the five prior instances hid.

SCOPE IS A FACT TOO, and it was the big one. Six gates decided what to look at with os.listdir - one
directory level. Every one was green AND CORRECT, because no template subdirectory exists today; every
one would have gone blind the moment anyone added templates/partials/, which is an ordinary act. A
single planted file carrying an emoji, a native confirm(), hand-rolled row markup, a dangling JS id
reference, a templated secret and an unregistered retrieval promise passed all six.

THE CONTROL IS WHAT MAKES THAT A MEASUREMENT: mojibake and docker-v already used os.walk, saw the
identical planted file, and convicted. So the cause was the listing, not the decoy.

COMMENTS ARE NOT CODE, AND COMMENTS ARE NOT CONTROLS. debug-routes matched `case subpath == "x"` in
raw text, so a case left in a commented-out block counted as a live handler - which is R-400's
original defect (seven dead controls on the page an operator opens when something is already wrong)
reached through the one door its own gate could not see. app-row-dedup's MUST_USE check had the same
shape: a commented-out {{template "app_list_row"}} satisfied it.

Stripping is deliberately crude in debug_route_gate, and that is correct there: its own docstring
insists on ten lines that cannot rot. A // inside a string literal truncates that line, which can
only ever HIDE a reference, never invent one - it fails in the safe direction.

NOT FIXED, and left open with its decoy rather than quietly patched: R-425, offbox-rename scans a
fixed three-entry FILES list, so banned NAS branding in a NEW offbox template passes. The scope was
correct when written and silently narrows every time the feature grows a file.

test_gate_decoys.py holds 10 decoys and declares COVERS, which felhom.eu's new decoy-coverage gate
AST-parses - a substring search for coverage would be the very shape this sweep exists to find.

No Go code. No version bump. No image. No golden owed.
Survey: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:39:17 +02:00
admin 3db62fc6b2 REPORT: the CI step is MEASURED, not assumed - job 481 shows the payload works
gates / gates (push) Successful in 12s
Gitea's act-runner does populate GITHUB_EVENT_PATH with a commits array carrying per-file lists.
Job 481 read 3 commits / 12 distinct paths, classified CODE (5 document, 7 code), and ran the gates
with --scope=code. Green.

Still not observed: the docs branch in CI, and an advisory in a CI log - the second additionally
needs a golden debt to exist at that moment. Neither is being arranged artificially.
2026-09-01 12:03:28 +02:00
admin 13dd00bcff gitignore Python bytecode - the new gate test writes controller/scripts/__pycache__
gates / gates (push) Successful in 13s
test_golden_notice.py imports controller_gates.py to assert the golden-notice is registered
non-blocking, which makes CPython write bytecode beside the scripts. felhom.eu already ignored it;
this repo did not, so it would have reappeared on every test run and been one careless 'git add'
away from being committed.
2026-09-01 12:01:54 +02:00
admin a4444088ad R-404: the golden NOTICE, in the repo where the debt is created - NOT A RELEASE
gates / gates (push) Successful in 12s
No version heading on purpose. No Go code, no image, no version bump; giving this one would create
the exact golden debt the change is about.

Until today this repo - where a release actually happens - had NO golden-currency check at all,
while felhom.eu ran one on every push including documents-only ones that can neither create the
debt nor clear it. The person who could act heard nothing; the person who could not act was
blocked, thirteen --no-verify uses' worth.

golden_notice.py is ADVISORY IN EVERY CASE, and that is the only correct behaviour rather than
timidity: at the moment a release is committed the golden legitimately does not exist yet, so
blocking there would refuse the commit that STARTS the process - and blocking later is the mistake
being undone.

NO SECOND IMPLEMENTATION: it IMPORTS felhom.eu/scripts/golden_currency_gate.py and calls that
gate's own released_versions()/newest_baked(), so it is the same comparison read in the other
direction. Cross-repo shape copied from instructions_gate.py; never a copy of the script, because a
copy recreates the drift these gates exist to detect. An absent sibling clone is INCONCLUSIVE and
silent about currency - it never guesses.

controller_gates.py GAINED A FIFTH `blocking` FIELD. It could not express a reporting-only gate at
all before: every registered gate's non-zero exit failed the run, so the only way to add a notice
was to give it the power to refuse a push. The capability was added rather than the notice
compromised (R-420). False for exactly one gate, and test_golden_notice.py asserts it stays one.

Tests N1-N4 with a positive control that every other gate is still blocking. RED-PROOF RUN: making
the debt branch return 1 fails N1 - in production that would refuse the commit that starts a
release.
2026-09-01 12:01:26 +02:00
admin a017367f9f REPORT: disclose the five red CI runs, the pre-push bypass, and two instruction defects
gates / gates (push) Successful in 12s
2026-09-01 10:58:06 +02:00
admin 4e13fdaaa0 REPORT: v0.232.0 - the determination, the walk's three extra findings, the golden
gates / gates (push) Successful in 13s
2026-09-01 10:50:52 +02:00
admin 62c6a8a98a docs(v0.232.0): CHANGELOG, three CONTEXT rulings, README (R-411/408/407, R-414, R-412a)
gates / gates (push) Successful in 14s
2026-09-01 10:36:52 +02:00
admin 8b55de734c R-414: the fallback scratch must also be DELETABLE - caught by live validation
gates / gates (push) Successful in 12s
The system-data fallback resolved a scratch fine and removeProofScratch then refused to
delete it: its accepted-roots list is built from REGISTERED drives, and a driveless box has
none. Observed on demo-felhom: 'refusing to remove ... it is not inside a proof root', with
the copy still on disk. Every nightly proof would have left one behind, growing forever, on
exactly the boxes the fallback exists for.

My defect, introduced with the fallback in the same session. The unit tests missed it
because every one of them registers a drive; the new pair deliberately does not, and the
second asserts the guard still REFUSES a path outside every proof root, so the fix is not a
widening into uselessness.
2026-09-01 10:29:08 +02:00
admin fcef8e069c one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:

  RestoreOffboxScratch      - the reported one
  OffboxRestorePrepareFull  - the SECOND request in the customer's own two-step full-restore
                              flow, and the one that actually shells `restic stats`. The UI
                              reaches it FIRST, so flagging only the restore would have left
                              the collision reachable by the ordinary path.
  RestoreSharesScratch      - R-411's exact shape on the shares tier: unlockStale + resticStep,
                              a live web caller, and its sibling PlaceSharesRestore has always
                              taken the flag.
  RestoreOffbox             - no production caller today, but the same dangerous pattern.
                              Flagged rather than left for a future caller to inherit.

OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.

THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.

R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.

R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.

The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.

And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.

R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.

16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
2026-09-01 10:19:36 +02:00
admin 9aea86cd48 REPORT: CI verdicts by run id - three controller runs green, felhom.eu 288 red on the declared golden debt only (12 of 13 gates OK)
gates / gates (push) Successful in 11s
2026-08-31 21:33:03 +02:00
admin 79f853fbbd docs: v0.231.0 REPORT + README + CONTEXT rulings + REUSE entries (R-87)
gates / gates (push) Successful in 12s
2026-08-31 21:31:04 +02:00
admin 303129e3af v0.231.0: the off-site proof gets a by-hand trigger, like its integrity sibling (R-87)
gates / gates (push) Successful in 12s
Without it the only way to see the job work is to wait for 05:30, which makes live
validation and any future diagnosis a next-day exercise. Same function as the scheduled
job - no second code path.

ONE deliberate difference from the integrity button: due-ness is NOT bypassed. There,
forcing means "check the store again", which is always answerable. Here due-ness IS the
target selection - an app is due when its newest snapshot has not been proved - so
ignoring it would mean inventing a second way to choose an app, exactly what having one
function prevents. When nothing is due the button says so, honestly.

Every other guard intact, including the single-writer flag: a hand-run during a backup
SKIPS exactly as the scheduled one would.

POST /api/debug/backup/offsite-proof, button beside "Restic integritas" on the debug page.
debug_route_gate pairs the two, so a button with no dispatch (R-400's shape) cannot ship.
2026-08-31 21:09:09 +02:00
admin e43b5ec07d v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
2026-08-31 20:55:34 +02:00
admin 2d802d75e8 docs: record the two documentation commit hashes in REPORT.md section 4
gates / gates (push) Successful in 12s
2026-08-31 14:39:41 +02:00
admin 1cfdde968f docs(v0.230.0): R-403 — CHANGELOG, CONTEXT rulings, README, REPORT
gates / gates (push) Failing after 13s
CHANGELOG v0.230.0, leading with the measurement rather than the fix: 120 082 104 B -> 7 036 B on
the shipped v0.229.0, reproduced before anything was built.

CONTEXT records three rulings: hollowness is a MANIFEST question and never a size question; the
guard fences one shape and NOT shrinking, because the derived-copy rebuild is a design decision; and
the rehydrate happens inside the restore because a follow-up job races the 5-minute capture. Plus
the shape the live run taught: a warning that fires on everything costs the same as the comforting
lie it replaces.

README documents the refusal, what each surface says, and why the capture job is deliberately not
guarded. REPORT leads with Part 1's result, carries the six red-proofs, the per-row Scenario D table
with its seven-app control, and eight observations including R-404 filed-not-acted-on and three
mistakes of mine recorded rather than tidied away.
2026-08-31 14:39:13 +02:00
admin b48a7fa326 R-403: the restore OUTCOME names the package's date, not the run's
gates / gates (push) Successful in 11s
Live on demo-hp the confirm said 11:43 (the preserved package) and the outcome said 14:23 (the
copy's newest run) for the same restore. A customer reading both cannot tell which one they had, and
one of the two is the flattering sentence. Part 2.3's rule is 'not a plain green success ANYWHERE',
and the outcome is an anywhere.

tier2UnitSourceMsg now asks UnitRestoreDate, the same resolver the confirm uses, so the two cannot
disagree. TestR403_OutcomeNamesThePackageDateNotTheRunDate pins it, with a negative control for the
ordinary case.
2026-08-31 14:32:29 +02:00
admin 5429d651ee R-403 fix, caught by the live run: the stale flag fired for every healthy app
gates / gates (push) Successful in 11s
UnitRestoreDate also compared the package's date against the run's and flagged 'older'. A recovery
unit is ALWAYS captured shortly before the run that mirrors it, so that comparison is true for every
healthy app. Measured on demo-hp: bookstack, kimai, opengist and privatebin all had src and dest
manifests at 12:03:49Z against a run at 12:14:24Z - perfectly healthy, and all four would have been
told their package was stale.

A warning that fires on everything is a warning nobody reads, which costs the same as the comforting
lie it was meant to replace. The second return is now UnitLegPreserved and nothing else.

TestR403_AHealthyAppIsNeverCalledStale pins it; red-proof: reinstate the comparison -> it fails.
2026-08-31 14:22:32 +02:00
admin 2358e561b7 R-403: a poorer copy must never delete a richer one
gates / gates (push) Successful in 11s
MEASURED FIRST, then fixed. On the shipped v0.229.0, on demo-hp, an app's Tier-2 copy went from
120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either) in ONE nightly
run, and the run recorded itself a success: 'Tier 2 copied docmost -> ... (14.9 KB, 0 leg(s), 0s)'.
Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/.

The mechanism was three individually-correct lines: RunTier2 guards the unit leg with os.Stat only
(does the folder exist), rsyncMirror is rsync -a --delete, and nothing between them compared source
to destination. An EMPTY unit is a folder that exists.

THE GUARD. One predicate, unitCarriesData/unitIsHollow (r403_hollow.go), asking the MANIFEST and
never the byte size - a big compose tree with no dumps is dangerous, a tiny unit for a tiny app is
fine. Fail closed on an absent or unparseable manifest. RunTier2 skips the unit leg when the source
is hollow AND the destination is not; the other legs still run, the run is not failed, and the skip
is recorded for the SURFACE (CrossDriveBackup.UnitLegSkipped + UnitPackageDate) as well as logged.

--delete STAYS and shrinking stays legal. 07 section 8 row 5's derived-copy rule is unchanged; the
fence is exactly one shape. TestR403_DataLegShrinkIsUnaffected is the guard on the guard.

THE HONESTY. A preserved package is older than the run that preserved it, so the card carries a
notice and the unit-restore confirm names the PACKAGE's date - read from the mirrored manifest's own
created_at, not from the status record - plus a clause saying why it is older.

THE CAUSE. RestoreTier2Unit now refills a hollow or absent primary unit from the mirror it just
restored from, INSIDE the call before returning. The hollow manifest was written two seconds after
a restore by the 5-minute capture job; any follow-up job races it. The capture itself is NOT guarded:
a capture describing an empty drive as empty is correct, and with the primary refilled there is no
hollow state left to describe. Never over a complete primary, never after a failed restore.

recordTier2Success and tier2UnitConfirmMsg keep their old signatures as thin callers, so no existing
test needed editing. New seam unitRehydrate, separate from tier2Mirror on purpose.

22 new Go tests. Red-proofs run and reverted: A6 (predicate -> size threshold), B1 (guard removed ->
the copy's 3 files are DELETED and the seam is called), B6 (a general never-shrink rule -> the shrink
case fails), C2 (only-when-hollow dropped -> the complete primary is overwritten).
2026-08-31 14:02:13 +02:00
admin fed272e62a docs: REPORT.md — delivery is DONE (golden 0.229.0 vouched, floor raised)
gates / gates (push) Successful in 11s
Section 6 now records the completed delivery instead of an owed bake; observation 2 records that
R-242's sixth conviction was paid the same day and the declared --no-verify bypass is historical.
New section 15 carries the bake, the round trip, the third independent reader, both proven
pre-gates, the three-field vouch, the R-120 gate passing, and the floor proven ACTING.
2026-08-31 12:38:15 +02:00
admin 0957f43d10 docs: record the two documentation commit hashes in REPORT.md section 3
gates / gates (push) Successful in 12s
2026-08-31 12:22:02 +02:00
admin 8aa95b5831 docs(v0.229.0): R-102 + R-103 — CHANGELOG, CONTEXT rulings, README, REUSE, REPORT
gates / gates (push) Failing after 13s
CHANGELOG v0.229.0. CONTEXT records three rulings: the source moves and the destination does not;
two predicates and not one wider one (R-356's cost restated); and a destructive operation reached
from a non-destructive surface carries the difference in the CONFIRM, not the label. README documents
the new action and route and corrects the coverage note to the measured count. REUSE maps the four
unit-directory-relative primitives and the new manager methods, with the traps.

REPORT covers the live drill on demo-hp (docmost, class B, primary unit moved aside — 3 volumes of 3
and 1 database of 1 in 28.65 s, accented filename byte-identical verified as hex, the app reading its
own row over TCP; Scenario D with the guest app.yaml also aside, secrets recovered=2/2), the settled
count (A=7 B=45 C=1, and why the earlier 9/43/1 was wrong), the five named red-proofs, and seven
observations including R-403 and a process error of mine that changed the box and is now in memory.
2026-08-31 12:21:28 +02:00
admin 4c8f0d2919 R-103: the Tier-2 refusal becomes an action
gates / gates (push) Successful in 12s
An app whose Tier-2 copy holds no file legs but a full recovery-unit mirror - 45 of the 53 catalog
templates - was told to press a button on a DIFFERENT page. Since R-102 the data it is asking for is
restorable from the copy it is looking at.

New POST /backup/tier2/unit-restore and backupTier2UnitRestoreHandler: same guards, same
restoreOpBlocked() refusal (R-351b), same async shape as the file restore beside it, plus a
fail-closed pre-flight so the app is never stopped for a mirror that could not be opened. The
outcome reuses unitRestoreOutcomeMsg and adds which copy overwrote the live data.

The row offers the action where the refusal was, in a danger style, as a SEPARATE button. The two
are not merged: one adds what is missing, the other overwrites. The confirm carries that difference
in words and names the copy's date - and says so differently when that date is only an ATTEMPT
(R-101). It is built from named Go constants rather than assembled inside an HTML attribute, so a
test can assert it verbatim; fmtTimeStr now delegates to a package-level fmtRFC3339Local so the
confirm and the outcome cannot render the same date two ways.

tier2NoCoverageMsg is NARROWED to the case that remains - no legs and no openable unit - and still
names the route that works. tier2UnitNotCoveredMsg is NOT deleted: it is appended where the FILE
restore ran and is still exactly true of it.

Tests C1-C2 and D1-D6 plus four more. Red-proofs: C1 (widen CanRestore to include HasUnit -> the
unit-only cases fail), D6 (drop EndRestoreOp from the handler goroutine -> 'the restore never
published a result').
2026-08-31 11:42:03 +02:00
admin 0f9b796615 R-102: the recovery unit on the second drive becomes a way back
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on
every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the
primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any
action in the product (07-backup-architecture 6.3, 7.2).

Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is
the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the
DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all
unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the
fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched.
reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now
named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second
one invented, which is what lets the acceptance test assert the volume leg's source directory.

Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a
parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken
inside RestoreFromRecoveryUnitAt, not beside it.

Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it
still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which
refused 40 running apps for months.

Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with
the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'.
Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the
primary -> fails with the mirror never reaching the redeploy, and with permission denied once the
primary tree is unreadable).
2026-08-31 11:30:34 +02:00
admin c732006d26 R-102 phase 1.1: split the recovery-unit path helpers, zero behaviour change
Every unit path helper took (nsRoot, stackName) and joined backups/primary/<stack>/... . That
hard-coded 'primary' is the mechanism of R-102: Tier-2 mirrors the whole unit directory to
<dest>/backups/secondary/<stack>/recovery-unit/ every night, and because no reader could NAME a
unit outside backups/primary/, that mirror has been captured for months and read by nothing.

Adds four unit-directory-relative primitives - UnitComposeDir, UnitManifestFile, UnitDBDumpDir,
UnitVolumeDumpDir - each taking the recovery-unit DIRECTORY itself. The four existing
(nsRoot, stackName) helpers become thin wrappers over them and keep their exact signatures and
their exact return values; every current caller compiles untouched.

ONE implementation, two callers - the rule restoreDockerVolumesFrom already follows in this repo.

TestR102_PathWrappersAreByteIdenticalToToday pins the wrappers against hand-written literals (not
re-derived from the helpers under test). Red-proof: UnitComposeDir join changed to 'compose2' ->
the test fails on all three fixtures.
2026-08-31 11:16:02 +02:00
admin 430fb4448d REPORT.md — golden 0.228.0 baked, vouched, floor raised; demo-felhom self-updated
gates / gates (push) Successful in 12s
2026-08-31 10:54:19 +02:00
admin 04abeb20c0 REPORT.md — correct observation 4: the memory was right, I did not follow it
gates / gates (push) Successful in 12s
2026-08-31 10:39:36 +02:00
admin 0989bb83c3 REPORT.md — fill in the felhom.eu commit hash
gates / gates (push) Successful in 12s
2026-08-31 10:39:13 +02:00
admin b670da748d REPORT.md — v0.228.0 (R-399 + R-400): proven live on demo-hp at both depths
gates / gates (push) Failing after 12s
2026-08-31 10:38:49 +02:00
admin 3c49dc8ea4 v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").

R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.

Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
2026-08-31 10:24:29 +02:00
admin 300d7e87d7 REPORT: R-359/R-397 shipped and validated; the measurement changed what it is worth
gates / gates (push) Failing after 12s
The headline is not the feature, it is what measuring it revealed: THE STRUCTURE
CHECK THAT SHIPS ON DOES NOT CATCH SILENT CORRUPTION. A pack corrupted without a
size change returned `no errors were found`, exit 0. Only --read-data caught it.
So R-399 is not merely a bandwidth question -- at the shipped default a class of
damage is not checked at all.

The three numbers R-399 needed are MEASURED, not estimated: store 134.3 MB / 67
snapshots; structure check 35.0 s; and the full curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none. Stated limit: they do not extrapolate.

Records four things that went wrong and were caught rather than shipped:
  - R-398 was MY OWN mistaken row. The seam already existed, Part 0 was not
    built, and building it would have HIDDEN the unlock --remove-all escalation
    from the assertions that must see it. Corrected, not closed.
  - the damage classifier matched restic's ORDINARY progress output; the
    negative control caught it (v0.227.1).
  - an exit code I misread through a pipe, corrected by re-measuring.
  - wire-contract convicted a field the hub cannot decode; allowlisted WITH A
    REASON rather than skipped, because building the hub display is a decision
    R-331 already ruled belongs to the operator.

And the sweep the task asked for: EIGHT debug buttons post to endpoints that do
not exist, not one. 24 referenced, 17 dispatched. Filed as R-400.

Not live-validated and each says why: the weekly firing (a week away), read-data
on a large store, and the join between "restic catches it" and "my code
classifies it" -- both proven, the join is not, and the seam is named.
2026-08-30 21:29:15 +02:00
admin 45770f2282 v0.227.1: the damage classifier matched restic's ordinary progress output
gates / gates (push) Successful in 11s
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and
re-pushing changed bytes under a live tag is the :latest hazard with extra steps.

looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A
HEALTHY restic check prints "check all packs" and "check snapshots, trees and
blobs" -- so any check that failed for a NON-damage reason, a connection dropped
mid-run for instance, would have been classified as a corrupted repository and
told the customer their backups may be damaged. That is the false alarm that
teaches an operator to ignore the true one.

Caught by the NEGATIVE control, built from the real bytes of a real passing
check on demo-hp. The spec made the negative control mandatory and this is what
it was for: a control that has only ever seen the failing case proves nothing.

Signatures are now phrases from restic's own error wording.

Also in this commit: CONTEXT.md records the three rulings (take the flag and
skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead
fields) plus the measurement a future session would otherwise assume wrongly --
THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job,
the route and the config, and corrects a line that listed four debug backup
routes when only two exist. REUSE gains three rows, including one that records
R-398 was my own mistake so nobody re-files it.
2026-08-30 21:22:23 +02:00
admin 0d52a42c17 R-359 + R-397: the off-site store gets checked, and the advertised check becomes real
gates / gates (push) Successful in 12s
Nothing ever verified that the off-site copies are still readable. The
whole-guest tier has verify jobs; the tier holding the customer's documents and
photos had none -- the complete set of restic verbs this controller used
contained no `check`. We would have found out at restore time, with a customer
waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain
`restic check`; we had never run it.

R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub
allowlists both event types and carries the Hungarian text for both, the
settings checkbox exists, and the debug button posts to /api/debug/backup/
integrity. Everything was built except the part that runs. SIXTH instance of
that shape in this project.

THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by
running `unlock --remove-all` and retrying, and its own comment records why that
is safe: every caller holds the in-process single-flight mutex, so any lock it
meets is stale. A check that did not take that flag could meet a LIVE prune's
lock from this same box, remove it, and retry over the top of it. So the check
TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly
backup behind it, and a skip costs nothing because due-ness makes tomorrow try
again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never
invoked, `unlock` never in any argv. Its red-proof prints the real thing --
restic running `check` while the flag was held.

DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful
check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape,
a dated check quietly missed and never caught up. No Weekly primitive added.

THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts.
"I could not look" is not "I looked and it is broken" -- R-339 already owns
reachability, and a second alarm for the same fact trains the operator to
discount the one alarm that means the backups are damaged. A timeout is
unreachable, never damage. A failure advances due-ness (a broken store must not
be re-checked nightly); a skip and an unreachable store do not.

Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by
design. A weekly success e-mail is how people stop reading their alerts.

The customer gets a SENTENCE; restic's words go to the log, truncated (R-379:
615 bytes of raw database text reached a customer once). read-data-subset ships
OFF and a malformed value is refused at read time rather than handed to restic,
where one typo would fail the whole check.

Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK --
those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero
still passes unmodified.

Also: the monitoring page stopped promising a Sunday job that never existed, and
the debug button got its dispatch case.

PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for
already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable
since the off-site tier shipped, and other tests drive restic-backed paths
through it. A resticStepFn seam would have been WORSE here -- it would replace
the `unlock --remove-all` escalation and hide it from the assertions that must
see it. R-358's AST ordering test is converted to a real execution test instead,
which immediately surfaced something the AST walk could not: unlockStale
legitimately runs before the restore.

Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages,
rc 0. All 12 controller gates OK.
2026-08-30 21:03:29 +02:00
admin e64c84aef8 REPORT: the golden debt named in section 12 was paid the same day
gates / gates (push) Successful in 12s
Golden 0.226.1 baked, published, round-trip verified, vouched, floor raised.
golden_currency_gate.py red -> green on the same command, so the three declared
--no-verify bypasses are historical rather than standing.

The fleet state changed with it: both demo machines run 0.226.1, and
demo-felhom got there by SELF-UPDATE rather than by hand -- which is the
positive observable that the floor is acting and not merely set.
2026-08-30 20:24:14 +02:00
admin c476ae51a5 REPORT: record the 0.226.1 patch and which version the live evidence describes
gates / gates (push) Successful in 12s
The section 6 live validation ran against 0.226.0, which was already on the box
when the fallback defect was found. Saying so rather than re-attributing the
evidence to 0.226.1: they differ only by CountsUnknown and its two tests, and
neither touches any path that validation exercised.
2026-08-30 19:53:21 +02:00
admin 36c13cfa0a v0.226.1: ship the CountsUnknown fix under a NEW tag, not a rebuilt 0.226.0
gates / gates (push) Successful in 11s
0.226.0 was already deployed to demo-hp by hand when the fallback defect was
found. Re-pushing a changed image under a tag that is already running somewhere
is the :latest hazard with extra steps -- two different images, one name, and no
way for a box to tell which it has. So the fix ships as a patch.
2026-08-30 19:51:52 +02:00
admin c0c8fe67bf An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero
result, which is honest" exposed that the sentence was FALSE.

A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's
fallback to RestoreApp -- which returns only an error, and whose signature is
deliberately out of scope -- would have printed "ez a mentes csak a
beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the
app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure
direction this whole change exists to remove, re-introduced by the change.

UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a
fourth sentence claiming only what is known -- the restore ran, the app is back,
and we cannot say what came back. RestoreApp's signature is untouched.

Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam
test was corrected too: its fixture has no recovery unit, so it exercises exactly
this path and had been asserting the wrong sentence -- it now asserts the
unknown, which is what pins the fallback to it.

IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate
written to stop findings dying in an overwritten REPORT.md caught a live defect
instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the
monitoring page advertises a weekly integrity check that does not exist) and
R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST
test) rather than leaving them in a file that is overwritten every session.

REPORT.md is the full run record: baselines re-confirmed, per-test results, the
five red-proofs with their observed output, the live validation with verbatim
Hungarian messages, what was NOT validated and why, teardown across three
layers, and the register 165 -> 167 -> 161.

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
2026-08-30 19:51:16 +02:00
admin e4e0aa8f46 REPORT: v0.226.0 shipped, three of four fixes proven live on demo-hp
Records what was validated and, in equal detail, what was not.

PROVEN LIVE (demo-hp, endpoints the UI invokes, evidence copied off the box):
  R-353  "A(z) opengist: 1 adatkotet visszaallitva -- az alkalmazas ujraindult."
         read off the customer's own wizard page, with real counts 1/1 volumes
         and 0/0 databases and correctly no database clause.
  R-360  refused in the exact flag state that produced the bug, and the planted
         canary file survived -- the consequence, not the branch.
  R-358  a mode=unit restore wrote {"schema":1,...,"full":false} at mode 0600
         with no .tmp left, and the gate logged place-to-live closed.

NOT live-validated, and each says why rather than being omitted:
  R-357  filling a real filesystem is a drill step, not a build step.
  R-353 Scenario B  NO app on demo-hp still has a data-less unit -- the spec
         named opengist from 21 August and it has since been recaptured (now
         1 volume dump). Manufacturing one means falsifying a manifest, which is
         the hand-set-state shortcut this project forbids.
  R-353 Scenario C and R-358's failed-download branch: unit-tested only.

Also recorded, because a near-miss that is quietly fixed teaches nobody: the
first B1 red-proof exposed a HOLLOW TEST OF MY OWN. With the gate removed the
run refused earlier, at the placement stat pre-pass, so `stops == 0` passed
against the pre-fix code. Fixture corrected and assertions reordered; only then
does the red-proof print THE APP WAS STOPPED (1 call(s)).

Register 165 -> 167 -> 161. Six rows compressed into CLOSED-ITEMS keeping title,
version, evidence and every sentence stating a rule; full original at
`git show e027b5d9`. No open row touched. ROADMAP not edited -- none of these
four ever had a row there, stated rather than silently skipped.

Teardown: this run provisioned nothing, across all three layers. Two throwaway
scripts and one canary directory were planted in guest 9201 and both removed.
2026-08-30 19:41:48 +02:00
admin b8af72764d R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21
backup-truth drill, all still in shipped code. They share one acceptance idea: a
restore surface must state what it actually did, and must refuse what it cannot
do.

VERSION NOTE. The task specifying this targeted v0.224.0 against baseline
f8c9390. Both were consumed earlier the same day by R-330 (0.224.0) and R-331
(0.225.0). Drift re-confirmed against live Gitea before the first edit, operator
authorised proceeding, every symbol the spec named re-verified present at the
real baseline e5eee50.

R-353 -- a restore that gave back nothing still said it worked.
RestoreFromRecoveryUnit returned only error, so the surface printed
"<app> visszaallitva (<snapshot>)." -- equally true of a run that returned an
entire dataset and one that returned nothing. The count already existed and was
discarded one line deep: restoreDockerVolumesFrom always returned it, the
wrapper threw it away. Now (UnitRestoreResult, error), carrying replayed counts
AND what the manifest LISTED, because zero-replayed has two causes that are
opposite news. Three cases, three sentences, and EVERY one is a claim about the
BACKUP, never about the app -- this path has no SafetyDump discriminator, and
07-backup-architecture 6.3 records that an absent dump says nothing about the
app (R-361 destroyed canonical .sql files for four months).

R-357 -- the destructive restore had no free-space gate. offbox_reconstitute.go
contained ZERO references to offboxFree; all three existing gates guard
non-destructive paths. The gate now sits before mapOffsiteRestorePaths,
writeSafetyDump and StopStack, so a refusal costs nothing. Position IS the fix,
which is why the test asserts StopStack was never called. No headroom multiplier
(matches PlaceOffsiteRestore; the x1.1 elsewhere predicts a download). Fail
closed on either probe <= 0 -- otherwise `free < need` with need==0 is FALSE and
an unmeasurable scratch sails through: a gate present and inert.

R-358 -- a failed download was offered as a good one. The gate answered "the
directory exists and is non-empty", which is exactly what a part-way restic run
leaves. Now a completion marker written 0600 atomically AFTER restic returns
nil, with any stale one cleared BEFORE it starts; both orders pinned by an AST
test because resticStep is not a seam. Both handlers refuse server-side: the
wizard flags control a button, and a hidden button is not a guard.

SCENARIO F ANSWERED, and worse than the question assumed: a unit-only scratch IS
reachable through the real flow, by the most ordinary route. "Ellenorzo
visszaallitas" (mode=unit, advertised non-destructive) writes the SAME directory
-- offboxRestoreScratchDir ignores `full` and --include limits what restic
extracts, never where -- so a customer who ran the SAFE restore was then offered
the destructive one over a unit-only copy. Filed R-396; the marker closes it.

R-360 -- the delete refused only while a BACKUP ran. IsRunning() is FALSE for the
whole of a verification restore; the five sibling handlers all use
restoreOpBlocked(). Its doc comment claimed it already did this, which is why
nobody looked -- corrected in place.

Red-proofs, each printing the pre-fix behaviour, in CHANGELOG and REPORT. The
first R-357 red-proof exposed a hollow test OF MY OWN and it is recorded rather
than quietly fixed: the fixture refused earlier at the placement stat pre-pass,
so `stops == 0` passed against the pre-fix code. Fixture corrected, assertions
reordered so a removed gate reports the outage rather than "no error returned".

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
2026-08-30 19:31:31 +02:00
admin e5eee501b5 R-331 (controller half): forward stats_known so the hub can tell empty from unmeasured (v0.225.0)
gates / gates (push) Successful in 12s
The hub's operator Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity
Unknown` for EVERY customer, because it rendered the report's `backup` object --
whose snapshot/size/integrity fields have had NO producer since disk-tier restic
moved to the host agent (slice 8C). buildBackupReport leaves them zero
deliberately and says so. Measured on demo-hp 2026-08-30 while that night's log
said `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`.

The live numbers were always in the report's `offsite` object, which the hub
already reads for its Offsite page and its fill/staleness alarms. The hub fix is
to render that -- and that made exactly ONE field mandatory that was not being
forwarded.

snapshot_count:0 means two opposite things: "holds nothing" and "never
measured". R-225 measured that confusion inside this repo (a rebuilt box
rendered 0 pillanatkep over a store really holding snapshot f3d9cd67), and
settings.OffboxTarget.StatsKnown fixed it for the controller's own UI. It was
never put on the wire, so the hub was free to make the identical mistake one
layer up -- and did. OffboxReportStatus.StatsKnown now carries it, omitempty, so
an older controller sends no key and a reader degrades to UNKNOWN, never to
EMPTY. Absence is ignorance, not emptiness.

The four dead BackupReport fields stay on the wire (historical reports in the
hub store must keep parsing) but now carry a warning naming R-331 and pointing
at Offsite. TestBackupReport_DeadFieldsStayZero fails the moment a producer
appears for one -- the prompt to update the hub card in the SAME change rather
than ship a field nothing renders.

RED-PROOF: drop `StatsKnown: t.StatsKnown` -> "a MEASURED empty repository
reported stats_known=<nil>". Tests assert the JSON the hub sees, not the Go
struct: measured-empty and never-measured must differ ON THE WIRE, which is the
entire point of the field.

Green gate clean: 28 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:38:10 +02:00
admin 45b52b6ed5 R-330: live validation on demo-hp — 3 scans inside the window, 0 events
gates / gates (push) Successful in 12s
v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven
through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks
stopped and restarted over 87s, three dead-app scans ran INSIDE that window
(16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps
that alarmed the night before on 0.223.0), zero app_start_failed pushed.

The scan count is the positive control, not decoration: an absent alarm is
equally consistent with "suppressed correctly" and "the scanner stopped".

A first run is discarded IN THE REPORT rather than quietly dropped -- it fired
52s after a controller restart, inside deadAppBootGrace (90s), where the scan
returns early and could not have alarmed whatever the code did. demo-felhom is
deployed but NOT independently proven and says so: its single app cycles in ~1s,
too fast for any 30s scan to land inside.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:12:26 +02:00
admin 92cebb8c95 R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.

The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.

The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.

It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
  1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
     every failure path, so the app alarms on the next scan;
  2. Begin REPLACES the set (one marker file = one operation);
  3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.

Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
  - drop markStopped from Begin      -> "suppressed at stop = map[]"
  - drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
  - pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.

Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 17:58:15 +02:00
admin f8c9390946 gate 11: register the observations gate, and mark up this repo's observations
gates / gates (push) Successful in 13s
The shared gate lives in felhom.eu/scripts/observations_gate.py and is invoked
across the workspace, exactly as reuse_refs_check.py and instructions_gate.py
already are. It is never copied.

REPORT.md's observations now carry their markers. Item 1 was the finding that
had no register row - only the first broken app per hour reaches the operator -
and it is now R-389. Item 4, the golden-bake runbook's missing `pveam update`,
is R-390. Items 3 and 5 are declared NOT-A-FINDING with their reasons. The
observations' text itself is unchanged; only the markers were added.
2026-08-23 13:53:18 +02:00