60f0a86bd4eb56cde7d6cfac584a61231e4d9c32
1058 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c0c8fe67bf |
An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero result, which is honest" exposed that the sentence was FALSE. A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's fallback to RestoreApp -- which returns only an error, and whose signature is deliberately out of scope -- would have printed "ez a mentes csak a beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure direction this whole change exists to remove, re-introduced by the change. UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a fourth sentence claiming only what is known -- the restore ran, the app is back, and we cannot say what came back. RestoreApp's signature is untouched. Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam test was corrected too: its fixture has no recovery unit, so it exercises exactly this path and had been asserting the wrong sentence -- it now asserts the unknown, which is what pins the fallback to it. IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate written to stop findings dying in an overwritten REPORT.md caught a live defect instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the monitoring page advertises a weekly integrity check that does not exist) and R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST test) rather than leaving them in a file that is overwritten every session. REPORT.md is the full run record: baselines re-confirmed, per-test results, the five red-proofs with their observed output, the live validation with verbatim Hungarian messages, what was NOT validated and why, teardown across three layers, and the register 165 -> 167 -> 161. Green gate clean: 28 packages, rc 0. All 12 controller gates OK. |
||
|
|
e4e0aa8f46 |
REPORT: v0.226.0 shipped, three of four fixes proven live on demo-hp
Records what was validated and, in equal detail, what was not.
PROVEN LIVE (demo-hp, endpoints the UI invokes, evidence copied off the box):
R-353 "A(z) opengist: 1 adatkotet visszaallitva -- az alkalmazas ujraindult."
read off the customer's own wizard page, with real counts 1/1 volumes
and 0/0 databases and correctly no database clause.
R-360 refused in the exact flag state that produced the bug, and the planted
canary file survived -- the consequence, not the branch.
R-358 a mode=unit restore wrote {"schema":1,...,"full":false} at mode 0600
with no .tmp left, and the gate logged place-to-live closed.
NOT live-validated, and each says why rather than being omitted:
R-357 filling a real filesystem is a drill step, not a build step.
R-353 Scenario B NO app on demo-hp still has a data-less unit -- the spec
named opengist from 21 August and it has since been recaptured (now
1 volume dump). Manufacturing one means falsifying a manifest, which is
the hand-set-state shortcut this project forbids.
R-353 Scenario C and R-358's failed-download branch: unit-tested only.
Also recorded, because a near-miss that is quietly fixed teaches nobody: the
first B1 red-proof exposed a HOLLOW TEST OF MY OWN. With the gate removed the
run refused earlier, at the placement stat pre-pass, so `stops == 0` passed
against the pre-fix code. Fixture corrected and assertions reordered; only then
does the red-proof print THE APP WAS STOPPED (1 call(s)).
Register 165 -> 167 -> 161. Six rows compressed into CLOSED-ITEMS keeping title,
version, evidence and every sentence stating a rule; full original at
`git show e027b5d9`. No open row touched. ROADMAP not edited -- none of these
four ever had a row there, stated rather than silently skipped.
Teardown: this run provisioned nothing, across all three layers. Two throwaway
scripts and one canary directory were planted in guest 9201 and both removed.
|
||
|
|
b8af72764d |
R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21 backup-truth drill, all still in shipped code. They share one acceptance idea: a restore surface must state what it actually did, and must refuse what it cannot do. VERSION NOTE. The task specifying this targeted v0.224.0 against baseline |
||
|
|
e5eee501b5 |
R-331 (controller half): forward stats_known so the hub can tell empty from unmeasured (v0.225.0)
gates / gates (push) Successful in 12s
The hub's operator Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` for EVERY customer, because it rendered the report's `backup` object -- whose snapshot/size/integrity fields have had NO producer since disk-tier restic moved to the host agent (slice 8C). buildBackupReport leaves them zero deliberately and says so. Measured on demo-hp 2026-08-30 while that night's log said `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`. The live numbers were always in the report's `offsite` object, which the hub already reads for its Offsite page and its fill/staleness alarms. The hub fix is to render that -- and that made exactly ONE field mandatory that was not being forwarded. snapshot_count:0 means two opposite things: "holds nothing" and "never measured". R-225 measured that confusion inside this repo (a rebuilt box rendered 0 pillanatkep over a store really holding snapshot f3d9cd67), and settings.OffboxTarget.StatsKnown fixed it for the controller's own UI. It was never put on the wire, so the hub was free to make the identical mistake one layer up -- and did. OffboxReportStatus.StatsKnown now carries it, omitempty, so an older controller sends no key and a reader degrades to UNKNOWN, never to EMPTY. Absence is ignorance, not emptiness. The four dead BackupReport fields stay on the wire (historical reports in the hub store must keep parsing) but now carry a warning naming R-331 and pointing at Offsite. TestBackupReport_DeadFieldsStayZero fails the moment a producer appears for one -- the prompt to update the hub card in the SAME change rather than ship a field nothing renders. RED-PROOF: drop `StatsKnown: t.StatsKnown` -> "a MEASURED empty repository reported stats_known=<nil>". Tests assert the JSON the hub sees, not the Go struct: measured-empty and never-measured must differ ON THE WIRE, which is the entire point of the field. Green gate clean: 28 packages, rc 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
45b52b6ed5 |
R-330: live validation on demo-hp — 3 scans inside the window, 0 events
gates / gates (push) Successful in 12s
v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks stopped and restarted over 87s, three dead-app scans ran INSIDE that window (16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps that alarmed the night before on 0.223.0), zero app_start_failed pushed. The scan count is the positive control, not decoration: an absent alarm is equally consistent with "suppressed correctly" and "the scanner stopped". A first run is discarded IN THE REPORT rather than quietly dropped -- it fired 52s after a controller restart, inside deadAppBootGrace (90s), where the scan returns early and could not have alarmed whatever the code did. demo-felhom is deployed but NOT independently proven and says so: its single app cycles in ~1s, too fast for any 30s scan to land inside. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
92cebb8c95 |
R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.
The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.
The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.
It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
every failure path, so the app alarms on the next scan;
2. Begin REPLACES the set (one marker file = one operation);
3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.
Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
- drop markStopped from Begin -> "suppressed at stop = map[]"
- drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
- pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.
Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
f8c9390946 |
gate 11: register the observations gate, and mark up this repo's observations
gates / gates (push) Successful in 13s
The shared gate lives in felhom.eu/scripts/observations_gate.py and is invoked across the workspace, exactly as reuse_refs_check.py and instructions_gate.py already are. It is never copied. REPORT.md's observations now carry their markers. Item 1 was the finding that had no register row - only the first broken app per hour reaches the operator - and it is now R-389. Item 4, the golden-bake runbook's missing `pveam update`, is R-390. Items 3 and 5 are declared NOT-A-FINDING with their reasons. The observations' text itself is unchanged; only the markers were added. |
||
|
|
1da2c9c6c6 |
docs(v0.223.0): REPORT, CONTEXT rulings, README severity contract
gates / gates (push) Successful in 11s
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate "warn" strings that are healthcheck statuses), the hub manifest's real location since the task's premise was wrong, all five red-proofs with the layer each guard sits at, the live walk in six steps with the hub's own records quoted, and the absent-intent count (0 of 8). Three things are reported that a tidier account would omit: red-proof 5 passed first time because the mutation was INERT; Scenario G was silently refused twice behind an HTTP 200; and the live Scenario A does NOT prove the customer gate, because demo-hp has no prefs row at all. CONTEXT records the severity vocabulary as a ruling with its mechanism, the intent ruling with its three-way handling of unknown, both fences, and two traps worth more than the fixes: a 200 can be a refusal, and a passing red-proof can mean an inert mutation. README: the event table said `app_start_failed | warn` - the defect, written down as if correct. Now `warning`, with the vocabulary contract and who receives what. `disk_critical` also corrected from `error` to `critical`, which is what fillwatch has always sent. |
||
|
|
9832760027 |
v0.223.0: the app-down alarm reached nobody (R-329), and the stop nobody heard (R-386)
gates / gates (push) Successful in 11s
R-329. NotifyAppStartFailures emitted severity "warn". The hub accepts exactly
{info, warning, error, critical} and silently coerces anything else to "info",
which severityNotifies then drops BEFORE both legs. Banner shown, event stored,
POST 200, no mail sent. One word.
This is the second time: DiskAlertKind.Severity emitted "warn" until v0.215.0
and its own comment records that every warning-level disk alert went to nobody.
A comment recorded the lesson and nothing enforced it. The guard is now an AST
walk over the whole controller - grep cannot work here, since "warn" appears
legitimately nine times as a healthcheck status vocabulary.
The sweep found exactly one bad severity. Its limits are stated: the walk cannot
follow a variable, so all six dynamic call sites are registered by name with the
values each can take, and a new one fails the test. Two of the six were found by
the guard, not by the hand sweep before it.
Also pinned: fillwatch.Band.Severity() returns "" for BandOK, which would vanish
the same way. It is unreachable because Check() notifies only on escalation -
but that safety lives in a different function from the one that looks unsafe, so
the test asserts the consequence rather than the mapping.
app_start_failed gains a customer toggle, DEFAULT OFF, per operator ruling. The
operator is mailed either way: processOperator never consults customer prefs.
It is deliberately NOT in operatorOnlyEvents, which would make the toggle a lie.
R-386. classifyRunStates decided "the customer stopped this" from the STATE, so
every stopped stack was assumed deliberate. Measured on demo-hp: privatebin
stopped out of band, nine scans, zero events, zero banner - while the comment
beside it claimed an out-of-band stop still alerts.
DesiredState already records the answer and has exactly one writer. Stopped ->
no alarm; Running -> alarm; absent -> UNKNOWN, keep today's behaviour AND say
so. Absent stays silent deliberately: reading it as "nobody asked" would email
about every app anyone ever stopped, fleet-wide, on the first cycle after
upgrade. The gap is bounded not silent - IntentUnknown is set and the names are
logged at INFO on the heartbeat cadence. failedRestart still lifts a Stopped
intent, or F-CRIT-1 re-opens. No new DesiredState writer.
Two settings toggles each governed two alarms. "Lemez figyelmeztetes (90%+)"
also wrote disk_critical, the drive-is-FAILING alarm. Now four honest toggles;
12 became 15. A no-op save stores the existing slice verbatim, so byte identity
is by construction - without that guard the defaults case reorders, which the
red-proof caught.
Test count 1504 -> 1522. Five red-proofs, five seen failing; one passed first
time and is reported - that mutation was inert, not the test weak.
|
||
|
|
14137efac5 |
docs(v0.222.0): REPORT, CONTEXT decisions, README state table + the ordering
gates / gates (push) Successful in 11s
REPORT.md overwritten with the full run: baselines and the hub's four numbers, the four red-proofs with the mutation and observed text for each, the five IsDownState consumers walked and named, the live walk in full with the old and new heartbeat lines quoted side by side, and the halt. CONTEXT records the decision - a dead supervised member is asked about before a failing healthcheck, because they are different questions and the second was answering the first - plus the fence that IsDownState did not move, the trap that three existing subtests pinned the defect, and R-386. README gains the `degraded` row, which the state table never had, and a note that the ORDER is load-bearing. Points at the new alarm-ladder architecture doc. |
||
|
|
5da11c4480 |
v0.222.0: ask whether a supervised member is DEAD before whether one is UNHEALTHY (R-384), and stop promising an undo copy nobody looked for (R-383)
gates / gates (push) Successful in 11s
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A two-container app whose database exits goes unhealthy BECAUSE it cannot reach that database - so the symptom the dead database causes was what suppressed the alarm for it. unhealthy is not a down state, so classifyRunStates never marked the app down and app_start_failed never fired. Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich failure, back through a different door. Two things moved, and either alone leaves the defect standing: the supervised test is hoisted above the unhealthy/starting/restarting returns, and "some members are up" now counts ANY member not in the down bucket. The old guard was running > 0, which made the R-51 block unreachable in exactly the case it was written for. IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy container is running and folding it in reintroduces the flapping that exclusion exists to stop. No new state was minted. Only the ORDER changed. The priority comment was rewritten because it asserted an ordering the code no longer has. Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they asserted an unhealthy/starting/restarting member beat an exited peer on unless-stopped, which pinned the defect as settled behaviour. They keep their intent with the down member given a benign policy. R-383. The double-failure message said the previous state's backup EXISTS, built from the returned path without asking the filesystem - and a missing file is one of the two ways that rollback fails. undoCopyPhrase now describes the copy from disk: present, partial, missing (still naming where it should be), or never written. Zero-length counts as missing. Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two halves of R-384 convict independently. |
||
|
|
da75603553 |
R-385 record: v0.221.1 gets its own CHANGELOG heading
gates / gates (push) Successful in 11s
0.221.1 was built, baked and vouched on 2026-08-23 with no entry of its own.
The prune-ordering fix (commit
|
||
|
|
f7881787f4 |
R-361 docs: CONTEXT decision, README, REPORT
gates / gates (push) Successful in 11s
Records the db_dumps decision with every consumer named, the trap that a stable db_dumps lets CaptureRecoveryUnit's already-current early return fire (so per-capture housekeeping must sit above it), and the NEGATIVE that a held app does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives it. |
||
|
|
810b18ab8e |
R-361 follow-on: the undo-copy prune must run above the already-current check
gates / gates (push) Successful in 11s
Excluding pre-restore-* from db_dumps made that list stable across restores, so CaptureRecoveryUnit's already-current early return began firing where it never had - and the prune, which sat after it, stopped running in exactly the case it exists for. Measured on demo-hp minutes after the change: four undo copies on disk against a cap of three. The prune is housekeeping on the dump directory and is independent of whether the manifest needs rewriting, so it belongs above the check. Pruning cannot disturb dbDumps, which no longer contains those names. Pinned by TestR361_UndoCapHoldsWhenTheUnitIsAlreadyCurrent; its red-proof moves the call back below the return and the cap fails at 5. |
||
|
|
968c968559 |
R-361: the safety dump destroyed the app's own database backup
gates / gates (push) Successful in 11s
writeSafetyDump called DumpOne into the app's OWN unit dir and renamed the result to pre-restore-* afterwards. DumpOne writes <stack>-<dbtype>.sql - the app's canonical dump - so every safety dump overwrote the app's real backup and then moved it away, leaving the app with no database backup until the next nightly run. A local restore-from-unit in that window tells the customer the app never had a database. The comment beside it asserted the rename meant it 'can never overwrite the app's real dump'. False as written, and believed for four months. Measured live before the fix: docmost and bookstack each held only pre-restore-* files and no canonical dump. DumpOneTo takes the final path and derives its own .tmp from it. DumpOne keeps its signature and calls it with the canonical name. writeSafetyDump asks for its own name directly; the rename is gone; the comment now states the invariant and how it is enforced. db_dumps no longer lists the undo copies. All three consumers of Manifest.DBDumps were grepped and named - all inside recovery_unit.go, none reads it for recovery. The files are neither deleted nor hidden. Tests 1485 -> 1493. FIVE red-proofs, TWO PASSED first time and both are reported: the behavioural tests inject the dump seam so a mutation inside DumpOneTo was invisible, and 1.3 had no test at all. Guards added at the layer each defect lives in; both mutations then convicted. |
||
|
|
2024ed9982 |
CONTEXT + README: the failure ladder and the operator ruling (v0.220.x)
gates / gates (push) Successful in 11s
|
||
|
|
1a2405e86f |
R-379: --clear-restore-hold now states the required restart
gates / gates (push) Successful in 11s
It runs as a second process: it clears settings.json but the running controller keeps its in-memory copy and goes on refusing. Measured on demo-hp - clear succeeded, file correct, start button still refused until a restart. Also records the lost-update window between the two processes, and why clearing through the running controller (the right shape) needs an operator tier the controller's HTTP surface does not have. |
||
|
|
5b52a5964d |
R-379 fix: the rollback must re-discover the DB container
gates / gates (push) Successful in 12s
Found by v0.220.0's own live walk on its first real run. writeSafetyDump
captures its DiscoveredDB before the stop; the DB-only start then re-creates the
container with a new id, so the rollback's docker exec hit a dead container and
sat in waitDBReady for 30s. The app was held for an infrastructure reason while
its data was recoverable.
Re-discover and match by {stack, engine} - what reimportDBDumpsFrom already did.
Fail closed when the container cannot be found.
No unit test caught it because they all inject the import seam and never look at
container identity. The new test asserts the identity handed to the import.
|
||
|
|
2c724c9283 |
R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts. |
||
|
|
0f3cf0dbb2 |
CONTEXT: R-356 design decision (v0.219.0)
gates / gates (push) Successful in 11s
|
||
|
|
c1dbb05ad6 |
README: where an off-site restore puts the data (R-356, v0.219.0)
gates / gates (push) Successful in 12s
|
||
|
|
cbc3fa589c |
docs: R-356 CONTEXT + REPORT (v0.219.0, proven live on demo-hp)
gates / gates (push) Successful in 11s
|
||
|
|
08eb1a6e3a |
R-356: the off-site restore refused every app that has no data drive
gates / gates (push) Failing after 12s
ReconstituteFromOffsite and PlaceOffsiteRestore both resolved the restore destination with the RAW HDD_PATH and read an empty answer as "the app is not installed". For 40 of the 53 catalogue apps that answer is correctly empty and permanent, so both actions refused forever for a running, healthy app — and told the customer to reinstall it "in the same place", which those apps never offer. Separate the two questions. "Installed?" is asked of ListDeployedStacks via a new Manager.isStackDeployed that fails CLOSED on a nil provider. "Where?" is answered by GetAppDrivePath — the same resolver CaptureRecoveryUnit wrote the snapshot with, so the restore aims at the place the backup came from. The 13 drive apps are unchanged: own drive, mismatch check, ack still required. A third refusal, with its own sentence, covers installed-but-no-resolvable-root. Fixtures that marked an app "installed" by giving it an HDD path now state deployment as its own fact. No assertion weakened. |
||
|
|
2da259af38 |
docs(v0.218.0): README names the volume leg and the compose-project attribution; CONTEXT records the blocker
gates / gates (push) Successful in 12s
The README's reconstitution sequence gains the volume replay it never had, and the AppBackup row states that DiscoverDatabases now prefers the compose project label. CONTEXT records the thing an operator most needs next: R-354's fix cannot reach the 40 apps that need it most until R-356 is closed, because the off-site restore still refuses outright for every app that declares no data drive. The live confirmation was therefore done on calibre-web and paperless-ngx. |
||
|
|
5ce3a44645 |
v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s
R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label, which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no -p). The old derivation stays as the fallback and an unresolvable attribution is now loud. Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller, not the catalogue. R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside the unit, so the off-site restore had no volume leg at all — proven live with planted files: calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking folder, and the restore reported success without it. For the 40 of 53 apps that declare no data drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay with an explicit directory: ONE implementation, two callers. Volumes replay before the database and inside the stopped window. VolumesReplayed reaches the message. The comment beside the skip was half false and is corrected; the half that still holds — the live unit is the local path's source — is named, and scenario D fingerprints the whole live unit across the operation. Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the code: scenario D passed with the unit guard removed because the fingerprint had been narrowed and was blind to the unit root. |
||
|
|
f94543ee5c |
v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok, go vet clean, -race clean on the changed package - all run and read BEFORE this commit. PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT installed in this case, so there is no own drive to ask) and reads manifest.json plus the captured compose/app.yaml. Local file reads only: no network, no restic, no restore. RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as "what your backup says" would be a fabricated fact. The prefill is labelled as coming from the backup and stays editable: a memory, not a lock. PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40 have none and their data goes to the system drive, which no screen said. Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for the 40-class a recorded placement is a fact to state, never a value written into a field that does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification. PART 4 - measured before theorising, on the live off-site target: snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds. The cause is the shape already on file, so the per-app size calls now run concurrently, BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted. OffsiteInventoryList had no test at all before this. TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a template doing index/eq against an undefined key errors at RENDER time: green build, green vet, green suite, 500 on the page. Four render tests, one per branch, because the existing deploy render test only renders AutoFields and never reaches these blocks. RED-PROOFS, mutation asserted applied then reverted to 0: A three template guards dropped (count asserted 3) -> the blank form returned P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory + what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first. NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit carries no db_dumps and no volume_dumps still reports a bare completion. |
||
|
|
985388c6e9 |
R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s
Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.
PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.
PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.
PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.
RED-PROOFS, each mutation asserted applied and reverted to 0:
B both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
C Mismatch forced false -> the silent divergent restore returned
E Known() forced true -> the fabricated empty prefill appeared
D Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.
Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.
NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
|
||
|
|
2fa1efc5e5 |
docs(REPORT): confirming cycle on v0.216.0, and persistence proven live
gates / gates (push) Successful in 9s
- 09:31:35Z on 0.216.0: '2 disk(s) evaluated, 0 alert(s)' — the count now matches the 2 persisted records, closing the disagreement that exposed R-335. - The 0.215.0 -> 0.216.0 redeploy replaced the container and the state file came back with a changed_at written by the PREVIOUS version, so the new container loaded the pre-restart record instead of re-baselining. Scenario L observed on real hardware, not just through the production-path unit test. - R-332 narrowed accordingly: what remains unproven is an already-ALERTED disk not re-alerting after a restart. |
||
|
|
330e4a051e |
docs(REPORT): v0.215.0 -> v0.216.0 run report
gates / gates (push) Successful in 9s
Includes the two clean live cycles, the warning-vs-warn notification_log proof, the 13 red-proof outcomes (A reported as a finding — the spec's mutation for it is not a valid red-proof), and section 14 on R-335, the aliasing defect found live in v0.215.0 and fixed in v0.216.0. |
||
|
|
90f2545679 |
fix(disk-health): one physical disk must be evaluated once per run (R-335)
gates / gates (push) Successful in 9s
Found on live hardware two hours after the v0.215.0 deploy, by noticing the release's own positive observable disagreed with its own persisted artefact: the check logged '3 disk(s) evaluated' while disk-health-state.json held two records. demo-hp's c11-scratch and felhom-backup are the same NVMe and share a durable id, so one disk was walked twice per run. Not cosmetic. The loop writes a disk's record before the next entry reads it, so the second copy of an aliased disk consumed the FIRST copy's write as its prior: the disk sustained against ITSELF and reached Hiba on a first sighting, defeating truth-table row 6 — the rule that separates a one-hour benign excursion from a false critical. It would also have emitted two identical events for one drive. Latent on demo-hp only because all counters are zero. Each diskKey is now evaluated once per run. Both entries stay marked seen so neither looks like a disappeared disk, and the card still renders both rows — the dedup is about state and alerts, not display. Red-proof run and reverted: deleting the guard makes the first sighting emit Kind:2 (Hiba-from-sectors) at 8 sectors. |
||
|
|
8144a70a72 |
docs(v0.215.0): CHANGELOG, CONTEXT decisions, README feature, REUSE entries
gates / gates (push) Successful in 10s
- CHANGELOG leads with the severity fix and the live warning-vs-warn proof. - CONTEXT records the settled decisions so they are not re-litigated: Hiba is the label for predicted failure (no fourth word); sustain before count and why; the provenance of 64/55/60; phase 2 owns the new SMART attributes because they are a wire change under G-1; phase 1 state is one record per disk, not a series. - README documents the 14-row ladder, the persisted state, the hourly cadence and the five message shapes. - REUSE pins the severity wire contract on PushEvent — the defect's real home, so the next typo'd severity is caught at the table rather than in production — and records priorFor vs cardPriorFor, which differ by one observation and make the chip disagree with the email if mixed up. |
||
|
|
34d83f5a02 |
feat(disk-health): poll hourly, not 6-hourly — measured, not assumed
gates / gates (push) Successful in 9s
Part 4 was gated on a measurement. On demo-hp (Tier 0) the controller's real /disks fetch — fetchDisks, the same path the check uses, not the 60s card cache — costs min 0.805s / median 0.821s / max 0.841s over 10 calls, all HTTP 200, across 3 physical disk rows (2 distinct devices). Median is 6x under the 5s bar, so the <5s branch applies and the interval drops 6h -> 1h. Why it matters: the one real failing drive's benign excursion lasted about ONE HOUR and cleared completely. A 6-hourly sampler can land either side of an excursion like that, see nothing, and then catch the terminal run half a day late. The smartd history that produced the whole analysis sampled every 30 minutes and only just resolved the shape. |
||
|
|
c24f1920d9 |
test(disk-health): Group L must run TWO checks after the restart
One check cannot distinguish a loaded state from a silent re-baseline — a forgetful controller is also silent on its first check. It betrays itself on the second, when the rebuilt prior makes the disk look newly sustained and it alerts all over again. Caught while building the companion red-proof: with the state load skipped, the single-check version still passed. |
||
|
|
bb50e1293c |
fix(disk-health): the alert that never sent — severity, a real Hiba level, and a memory that survives a restart
Three defects made the disk-health feature silent in exactly the case it
exists for. Evidence: felhom.eu documentation/audits/DIAG-smart-passed-trap-2026-08-14.md
1. SEVERITY (the one that changes whether anything arrives at all).
NotifyDiskHealthDegraded emitted severity "warn", which is NOT in the
hub's accepted set {info,warning,error,critical}. The hub coerced it to
"info" (hub/internal/api/handler.go) and severityNotifies dropped it
(hub/internal/notify/dispatcher.go), so every Figyelmeztetes-level disk
alert was filed as an informational notice and emailed to NOBODY, on the
customer and the operator leg alike. Now "warning". DiskAlertKind.Severity()
is exported so the contract is checkable from any package.
2. NO LEVEL ABOVE "worth an eye". smart_status.passed CANNOT fail on
unreadable sectors (attrs 187/197/198 all carry thresh 0 and a normalized
value floors at 1), so Hiba was unreachable for this whole fault class.
DiskVerdictFor now takes a DiskPrior and implements a 14-row top-down
ladder: sustained unreadable sectors, a count too large to be a blip (64),
unreadable+remapping together, overheating, NVMe critical flag or spent
endurance all reach Hiba. No fourth label — predicted failure is "Hiba".
3. IT SPOKE ONCE, AND FORGOT ON RESTART. The baseline was in-memory, so a box
that rebooted while a disk was failing never alerted again; and between 8
and 352 sectors nothing was emitted at all. State is now persisted
(disk-health-state.json, atomic tmp+rename), the decision compares against
the last ALERTED verdict (collapsing flaps to one alert while letting a
genuine escalation fire immediately), and a disk already at Hiba re-alerts
once it has BOTH doubled its count and waited out a 24h cooldown.
The card replays the same prior the check used (diskRecord.PriorSawUncorrectable)
so the chip and the email cannot disagree — the property the shared verdict
function exists to guarantee, now pinned rather than asserted.
Tests: 12 scenario groups A-L. Group L builds the Server through web.NewServer,
the same call main.go makes, over a real file.
|
||
|
|
3e3ee94b7b |
REPORT: live 422 proven on hardware; R-308 withdrawn (my quoting bug, not a stale credential)
gates / gates (push) Successful in 12s
|
||
|
|
ae10f64806 |
REPORT: controller v0.214.0 — the screen stops hedging, and the claim guard grew a surface
gates / gates (push) Successful in 18s
|
||
|
|
3ed5e3e770 |
v0.214.0 — the recovery screen stops hedging about a code it can now check (R-311)
gates / gates (push) Successful in 13s
MinAgent: 0.129.0 What was already right: the screen did not bluntly accuse. R-222/R-226 hedged, naming both causes and the kept package, and saying it could not tell them apart. That was honest - and it could not tell them apart because nothing ever looked. Agent v0.129.0 looks, so the hedge becomes an answer. New class RecoveryCodeOpensRetained on HTTP 422, gated by FeatureRetainedRecoveryClass (MinAgent 0.129.0). The gate is SEPARATE from the R-224 one because the two name different agent versions and a box can sit between them, where a 422 is a shape we did not design. ClassifyRecoveryFailure therefore takes both flags; the compiler found every call site. The message says the code is correct, names the supersession date, says the earlier package is kept, and says the CURRENT backups are unaffected - the half a customer will otherwise assume wrong. It promises NO restore: there is no in-product route to a set-aside store (R-312) and the retained package may itself predate the repository-password field. It routes to support, which can do it. The claim guard grew a surface and immediately convicted something. It scanned templates only, while every recovery message is a Go string in a handler - the highest-stakes copy in the product, never scanned. It now scans recovery_handlers.go too, and found a PRE-EXISTING unregistered claim on its first run. Six handler tests asserting which SENTENCE the customer sees; red-proofs asserted applied, including: 422 unconditional makes an agent that never looked read as having looked, and routing 400 to the new class congratulates a mistype. |
||
|
|
3168a78935 |
REPORT: v0.213.0 pinned-fingerprint condition, red-proofs, claim guard
gates / gates (push) Successful in 12s
|
||
|
|
89712563a0 |
R-302: the abandon banner promises only what the box can still see is true
gates / gates (push) Successful in 10s
The retrieval clause rendered unconditionally on every page and is false on a reachable state - the same screen where the orphan card says we cannot tell. The condition is a fingerprint PINNED at the decision, not a comparison against the current key. The obvious proxy asks about the wrong key: the set-aside copies were written under an older key the box no longer has, so on a twice-rebuilt box the proxy promises about copies nothing can open. Demonstrated - under the proxy, the replaced-package and legacy cases both flip back to promising. The pin is a recorded assumption and says so: nothing on the box records which key wrote those copies. Empty is not a match. A countdown started before this carries no pin and takes the cautious branch, not a backfill. A sweep of all 36 templates found a fourth instance (backups page, same condition applied) and a fifth (the confirmation screen, correctly left alone - true at the moment of the decision). New retrieval_promise_gate registers each claim with a reason rather than banning a verb: a string ban failed twice, and the honest replacement copy contains the stem. |
||
|
|
1b66010298 |
REPORT: v0.212.0 orphan card second promise
gates / gates (push) Successful in 16s
|
||
|
|
68f3e12398 |
R-299: the orphan card's second promise, and a guard that matched one inflection
gates / gates (push) Successful in 14s
The explanation paragraph - the always-visible half of the card - still ended 'a hozzajuk tartozo helyreallitasi koddal kesobb visszaallithatok lehetnek', the same unevaluable claim v0.211.0 removed from the confirm block below it. It survived because the spec called that line accurate, and because the regression guard asserted the SINGULAR form while the card carried the plural, which does not contain that substring. The guard now matches the stem, so any conjugation fails it. The two accurate halves are kept. Also: the guard's failure message sliced rendered HTML at a byte offset and cut Hungarian mid-character; it now slices on rune boundaries. |
||
|
|
f87be3575f |
REPORT: correct installer publication status
gates / gates (push) Successful in 12s
|
||
|
|
38f4535bfa |
REPORT: golden 0.211.0 baked and published; only the Day-0 vouch remains
gates / gates (push) Successful in 14s
|
||
|
|
397d62136f |
REPORT: v0.211.0 written, not delivered - bake and Day-0 approval outstanding
gates / gates (push) Successful in 13s
|
||
|
|
86a78c6767 |
R-294/R-295: orphan card stops promising restorability; one name per secret
gates / gates (push) Successful in 14s
The orphan card told a customer their set-aside off-site history may be restorable later with their recovery code. The discriminator lives on the hub and no wire field carries it, so the box rendering that card cannot evaluate the promise. Copy replaced per the spec: state what happens, decline what we cannot know and say why, name a route. The claim page called the same three-word dashboard code two different names depending on branch, one of which collides with the ten-word escrow code. Retired 'Visszaallito kod'; the name is now constant and the sentence changes. Naming only - a test pins that a reset code is still accepted. secret_in_markup_gate no longer convicts Go template comments, which are stripped before render; still convicts a real rendered secret. |
||
|
|
b762a37097 |
R-280: attach list from mounted-but-unregistered filesystems; two-clicks promise made conditional
gates / gates (push) Successful in 17s
After a reinstall the data drive could not be re-attached through any dashboard route: both candidate lists came from the agent's unclaimed-disk scan, and the rebuilt box's drives are claimed. The restore page said it was two clicks while pointing at an empty picker. The attach list now also carries the controller's own mounted-but-unregistered filesystems. initialize is untouched, so the format wizard's system/backup protection is unchanged. The 'two clicks' sentence is conditional on the picker being non-empty, and says something true and actionable when it is not. |
||
|
|
c732fe1283 |
v0.210.0 — R-259 and R-258: two pictures that were not true
gates / gates (push) Successful in 18s
Both are one shape: something the box already knows, drawn as its opposite. R-259 — A DISK WE FAILED TO READ WAS DRAWN AS A HEALTHY EMPTY DISK. readDiskUsage (internal/system/info_linux.go) logged a statfs failure at DEBUG and returned, leaving the caller's TotalGB/UsedGB/AvailGB/Percent at zero — and usageColor(0) is "nominal". The dashboard's most-looked-at meter therefore rendered "0.0 GB / 0.0 GB (0%)" with a 0%-wide bar in the healthy colour. "We could not look" and "there is plenty of room" were the same picture. readDiskUsage now returns whether the measurement succeeded; SystemInfo gains DiskKnown and HDDKnown (HDDConfigured is not a substitute: it says a path was configured, not that reading it worked); and the template draws NO figure, NO percentage and NO meter fill when unknown, saying "A tarhely merete most nem olvashato ki." instead. A healthy box is byte-identical, colour band included. This session rules the convention (felhom.eu CONTEXT.md S-39): an explicit `...Known bool` companion beside the figures, checked in the template — the shape Offbox.StatsKnown already uses, whose own comment says "a 0%-wide bar over an unread store is a picture of emptiness, and a picture is a claim". Pointers and separate error fields are both legitimate Go, but a codebase with three dialects cannot be gated (ROADMAP G-3 was blocked on exactly this). Existing call sites NOT converted. R-258 — THE PER-APP BACKUP TICK WAS GREEN ON PRESENCE, AND RED ONLY ON A GLOBAL CONDITION. buildAppBackupRows set Tier1LastStatus from status.LastDBDump.Success, which is the box's single most recent dump RUN, whichever app it belonged to. An app whose own dump failed showed a tick as long as some other app dumped successfully afterwards; an app with no database took the nil branch and went green on the mere existence of a restore point. appDumpVerdict now reads THIS app's own entries in DBDumpStatus.Results (matched on DumpResult.DB.StackName, failure = non-nil Error). Three states: any failing database -> error; all clean -> ok; no result recorded -> NO verdict and no icon, titled "Errol a mentesrol nincs eredmenyunk." The recovery unit carries no per-run outcome of its own, so green cannot honestly be derived from presence. The global tier1DBStatus label is untouched — it is correct as a global. RECENCY IS DELIBERATELY NOT ADDED. A tick over a three-week-old restore point is a real weakness, but an age threshold means inventing a number and the time is already printed beside the icon. Recorded as an observation, not changed. AN EXISTING TEST WAS ASSERTING THE DEFECT AND WAS CORRECTED, NOT DELETED: TestBuildAppBackupRows_Tier1FromRestorePoints expected "ok" for a status with no LastDBDump at all — green from nothing but a file's existence. It now expects no verdict; its real subject, the Tier1LastRun time, is unchanged. The dashboard test EXTRACTS the meter block from the shipped template rather than copying it: a copied block drifts, and a drifted copy passes while the page it claims to cover has changed — the fixture-is-not-the-wire mistake this project has now hit twice. Six red-proofs across both parts, each with the mutation asserted applied. No new tag on any declared wire — report/builder.go maps into its own types and is untouched; wire_contract_gate.py confirmed green. go build / go vet / go test ./... green (28 packages), controller_gates --fast all OK, both run separately from this commit. |
||
|
|
fcffaf573a |
v0.209.0 — R-247: the box stops saying a false thing about its own recovery package
gates / gates (push) Successful in 17s
The answer was on the wire and was discarded at the boundary, for the third time. The hub has sent `escrow_stale` in the report ACK since v0.57.0 (json:"escrow_stale,omitempty"). report.EscrowStatus had no field for it, so encoding/json dropped it, and an empty restic_pw_sha256 had exactly one possible reading here: "hash-less supersession". On demo-hp that reading was false in EVERY clause for four days, and the box told the customer so in its own words. The hub HAD the hash and was withholding it because the escrow row carries a stale flag (R-246); there had been no supersession; and the bundle DID cover the password — the hashes matched exactly. Fixed by receiving the field. EscrowStatus.Stale decodes, and reconcileEscrowed tells the two conditions apart: a withheld hash now reports that the hub has flagged the row and is withholding, that this box therefore cannot verify its bundle either way, and that it is NOT established that the bundle fails to cover the password. The genuinely hash-less case keeps its wording. Deliberately NOT changed, and said rather than skipped: the stale verdict itself (the hub's flag is still the hub's verdict; runs still continue), and the customer-facing Hungarian card copy. Clearing the wrong flag is an operator act hub-side (R-246); re-wording the card is UI work with its own review path. This change is the wire and the diagnosis. Found by felhom.eu/scripts/wire_contract_gate.py (G-1), which was built first and seen failing on 40 fields before anything was fixed, and which now refuses any new field of this shape. go build / go vet / go test ./... green, run separately from this commit. |
||
|
|
37b5ba08a7 |
REPORT: v0.208.0 — both R-254 sites, the guard's measured holes, and what the live read could not prove
gates / gates (push) Successful in 13s
Records the deliverables, and is explicit about the limit on the live half: the curl of an app info page could not be done, and names exactly what was tried — crafty-controller is the only app declaring initial_credentials and is deployed nowhere, and demo-hp's dashboard password in ~/.config/credentials no longer authenticates (200 with no session cookie). A probe of the new routes was discarded because its control killed it: real and bogus paths both 302 behind the auth middleware. §7.2's answer including the part that contradicts the task's premise: no line in the repo says 'no silent auto-fill'; the rule is CONTEXT.md:2070 about accidental EMPTY-password deployments. The hidden input is deliberate and untouched. §7.4's measurement: the gate covers all 36 templates and catches a launder through a local variable, but is blind to a secret under a neutral page-data key — the exact shape of site two. Runtime coverage is 4 of 27 pages. Filed as R-255 rather than described as complete. §7.3: no evidence of actual exposure on the fleet, with the limit stated — it is a current-state measurement and nothing recorded reads, which was part of the fault. Also corrects v0.207.0's report: html/template STRIPS HTML comments; they do not ship in the response body. Measured. |
||
|
|
27d1165962 |
v0.208.0 — R-254: the last two secrets leave the page source, plus a gate against a fourth
gates / gates (push) Successful in 17s
Site one. app_info.html rendered {{.InitialCreds.Password}} into a hidden span —
a REAL per-install credential, read live out of the running container, in the
response body of every render. The page now carries the non-secret half plus a
boolean; the value comes from POST /apps/<slug>/initial-credentials/reveal, which
RE-READS the container rather than serving a cached copy (caching it in the
handler would put it back in the body one layer in). no-store, CSRF-covered,
logged as an act. Both buttons go through it. A reveal that cannot read the value
SAYS SO rather than returning an empty string that renders as a blank password.
Site two, established before changing. The hidden input is NOT the defect and was
left alone: it fires only pre-deploy, and README §318 documents why the value must
round-trip — the customer notes the generated secrets down and submitting them
back is what makes the saved value the same one they saw. The defect was the
neighbouring READONLY input, which on an ALREADY-DEPLOYED app rendered the secret
into a page with nothing to submit. Fixed by POST /stacks/<name>/auto-field/reveal,
authorised by requiring a type:secret auto-field of that stack. Both directions
pinned.
The premise that this contradicted a repo rule does not hold: the rule is
CONTEXT.md:2070 'Password fields require explicit input — prevents accidental
empty-password deployments', about EMPTINESS. No line in the repo says 'no silent
auto-fill'.
The gate. scripts/secret_in_markup_gate.py, registered in controller_gates.py,
convicts any template expression that names a secret unless allowlisted with a
reason. Its limits are MEASURED and in its docstring: it catches a launder through
a local variable (the assignment names the secret) but is blind to a secret
arriving under a neutral page-data key — verified both ways. That is the shape of
site two, which this gate would NOT have caught. The runtime body assertion covers
all shapes but only 4 of 27 page templates; the other 23 are R-255, filed rather
than glossed. Two nets, different holes, both named.
Correction to v0.207.0's report: HTML comments do NOT ship in the response body
here — html/template strips them, text/template does not. Measured. A red-proof
planting a secret in a comment therefore correctly does not fail.
|