From c297b9f85e34344d0c21df54d1c61766676ffcfb Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 22 Aug 2026 13:26:10 +0200 Subject: [PATCH] R-356 docs: correct R-107 in the architecture, record the design, refresh STATUS, compress the register 07-backup-architecture.md: three places said no offsite action unpacks the named-volume tars. R-107 closed in controller v0.218.0; all three corrected with a dated [FACT], the old sentence kept in the past tense. R-102 is NOT closed and the correction says so explicitly. New [DESIGN] paragraph in 6.3: the restore destination is resolved by the same rule as the capture destination, and the wrong-disk refusal applies to apps that have a drive to get wrong. Carries the 13/40 measurement. STATUS.md was internally contradictory - nothing waiting, and one decision waiting, for something the same page recorded as shipped. 218 -> 102 lines; the deciding section now says what happens if nothing is done. R-356 compressed into CLOSED-ITEMS.md; OPEN-ITEMS 327109 -> 325236 bytes. Drill record and 16 evidence files for the live walk on demo-hp. --- REPORT.md | 343 ++---------- STATUS.md | 270 +++------- .../architecture/07-backup-architecture.md | 50 +- .../README.md | 78 +++ .../evidence/01-planted-manifest.sha256 | 15 + .../evidence/02-offsite-runs.txt | 18 + .../evidence/03-destruction.txt | 17 + .../evidence/04-prepare-scratch.txt | 3 + .../evidence/05-scratch-contents.txt | 5 + .../evidence/06-reconstitute-post.txt | 3 + .../evidence/07-reconstitute-log.txt | 30 ++ .../evidence/08-restored-manifest.sha256 | 15 + .../evidence/09-outcome-message-verbatim.txt | 7 + .../evidence/10-controller-log-phase1.txt | 400 ++++++++++++++ .../evidence/11-13class-planted.sha256 | 3 + .../evidence/12-13class-destruction.txt | 13 + .../evidence/13-13class-scratch.txt | 20 + .../evidence/14-13class-restored.sha256 | 3 + .../evidence/15-13class-outcome-message.txt | 7 + .../evidence/16-controller-log-phase2.txt | 500 ++++++++++++++++++ documentation/backlog/CLOSED-ITEMS.md | 1 + documentation/backlog/OPEN-ITEMS.md | 1 - 22 files changed, 1316 insertions(+), 486 deletions(-) create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/README.md create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/01-planted-manifest.sha256 create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/02-offsite-runs.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/03-destruction.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/04-prepare-scratch.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/05-scratch-contents.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/06-reconstitute-post.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/07-reconstitute-log.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/08-restored-manifest.sha256 create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/09-outcome-message-verbatim.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/10-controller-log-phase1.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/11-13class-planted.sha256 create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/12-13class-destruction.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/13-13class-scratch.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/14-13class-restored.sha256 create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/15-13class-outcome-message.txt create mode 100644 documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/16-controller-log-phase2.txt diff --git a/REPORT.md b/REPORT.md index 066eed68..319ef548 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,299 +1,72 @@ -# REPORT — one register, and a rule that keeps it readable (2026-08-22) +# REPORT — R-356 doc corrections, register housekeeping, and one gate fix (2026-08-22) -**Records and process only. No controller, agent or hub code. No release, no bake, no machine -contacted.** Previous report preserved at -`documentation/audits/REPORT-mis-called-defect-and-sweep-2026-08-22.md`. +Companion to `felhom-controller` v0.219.0 (R-356). This repo carried the architecture correction, the +drill record, the register move and one genuine gate defect found on the way. ---- +## 1. `documentation/architecture/07-backup-architecture.md` — R-107 was closed and the doc said otherwise -## 1. Baselines — confirmed +Three places, each corrected with a dated **[FACT]** citing `offbox_reconstitute.go` `volReplay` and +controller v0.218.0: -Controller **0.218.0**, agent **0.130.0**, golden **0.218.0 vouched**, floor **0.218.0** — all read -back from the hub, not assumed. Register ceiling was **R-375** (grepped, not trusted). All four repos -clean, `HEAD == origin/main`. +- §6.3, the Tier-3 row — was *"no offsite action unpacks the named-volume tars it captures"*. +- §8, matrix row 4 — was *"the volume tars in either copy are unreachable"*. +- The R-107 index row. -**The prompt's measurements, re-measured.** `ROADMAP.md` 249 lines / 239,306 B and `CONTEXT.md` -2,580 lines matched **exactly**. `OPEN-ITEMS.md` read 691 lines / **672,376 B / 286 entries, 134 -closed** against the prompt's 642,731 / 272 / 104 — the difference is **this project's own last -session**, which added 17 rows, plus a wider closed-vocabulary on my side (I count `RESOLVED` and -`FIXED` as terminal). Longest single entry **16,108 B (R-193)**, not 15,965 — same entry, since grown. +**The old sentence's history is kept, not deleted:** each correction says what was true, until when, +and what closed it. A correction that erases what was believed leaves the next reader no way to tell +a fixed gap from one that was never noticed. ---- +**R-102 — the Tier-2 half — is NOT closed, and the correction says so explicitly** so it cannot be +read as covering both. The Tier-2 row stands exactly as written. -## 2. The 59-versus-72 reconciliation — **both were right when taken, and neither should size anything** +## 2. §6.3 — the R-356 reasoning recorded as reasoning, not as a closed row -| method | value | stable? | +A new **[DESIGN]** paragraph: the restore destination is resolved by the same rule as the capture +destination (drive if the app has one, system data path otherwise); the refusal that protects a drive +app from being restored onto the wrong disk applies to apps that **have a drive to get wrong**. It +carries the 13/40 measurement and points at `felhom-controller/CONTEXT.md`. + +## 3. `STATUS.md` — the contradiction is gone, and the page is one screen + +It said nothing was waiting while also saying one decision was waiting — to publish agent 0.130.0 — +which the same page recorded as already published (R-347, closed). **218 → 102 lines.** + +"Waiting on you" now lists **four** real items, and — per the exemption — **each says what happens if +the operator does nothing**. The two new ones are this release's hand-off: vouch a golden carrying +controller 0.219.0, then raise the floor last, in a separate save. + +## 4. `documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/` + +The live walk on `demo-hp`: both classes, both messages verbatim with their byte counts, both accented +filenames as explicit hex, and 16 evidence files. **Copied off at the end of each phase.** Nothing was +reverted during this drill, so no intermediate teardown could have taken it. + +## 5. Register housekeeping (N.7) + +R-356 compressed out of `OPEN-ITEMS.md` into `CLOSED-ITEMS.md`, keeping its title, shipping version, +evidence path and every sentence stating a rule, plus the pointer +`git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` for the full original text. + +| file | before | after | |---|---|---| -| ids **mentioned** in ROADMAP minus ids mentioned in the register | **72** at commit `5a7502b`, **61** today | **no** | -| ROADMAP **rows** minus register **rows** | **90**, unchanged across all five commits checked | **yes** | -| …of those 90, marked shipped/closed/killed/ruled | **59** | — | -| …of those 90, **not** so marked | **31** | — | +| `OPEN-ITEMS.md` | 327 109 bytes | **325 236 bytes** | +| `CLOSED-ITEMS.md` | 61 580 bytes | **63 507 bytes** | -**Why 72 became 61 without anything moving:** my own **R-369** row, written last session, *names* 11 -of those identifiers in its prose. A "mentioned" count therefore falls when someone merely writes -about the problem. **That measure is unusable for sizing a migration.** +## 6. A real gate defect, found by CI going red -**And 59 is real** — it is the *history* half of the row-based population. So the prompt's figure and -mine were measuring the two halves of the same 90. **The count that matters is 31**, and the gate -later found the true figure is **32** (§4). +CI run **387** (job 386) failed `instructions_gate` on commit `08eb1a6` while run 73 on its parent had +passed. The cause was not the push: `ef6ac6f` (this repo, the same day) compressed closed rows out of +`OPEN-ITEMS.md` into `CLOSED-ITEMS.md`, and `register_state()` read only `OPEN-ITEMS.md` and +`ROADMAP.md`. **Every citation of a compressed item became "a reference to nothing"**, failing the +next push in a sibling repo for a rule file nobody had touched — and it would have fired again on this +task's own R-356 compression. ---- +`CLOSED-ITEMS.md` is now the third source. It answers *does this ID exist* and answers *closed* for +the rows it owns; `OPEN-ITEMS.md` remains the sole authority on **openness**, so a row it claims as +open is not overridden. **Both controls still convict:** an ID present nowhere fails, and a citation +claiming a closed item is still open fails — each watched failing, then watched clearing. -## 3. The population, and the rule used to sort it +## Gate state -**The rule, stated so it need not be re-invented:** - -> **Does the item assert something about the shipped product that a reader could go and check, and -> find false?** If yes it is a **finding** and belongs in the register. If it proposes something that -> does not exist yet — a feature, a spike, a curation task — there is nothing to be wrong about, and -> it stays in the roadmap as an intention. - -Two refinements the population forced: **an owed operator decision moves too** (it is work owed, not -an idea), and **an item already satisfied moves as CLOSED**, so nobody redoes it. - -| group | count | disposition | -|---|---|---| -| **open work — findings and owed items** | **17** | **moved to the register**, keeping identifier, evidence and original filing date | -| **ideas and proposals that were never findings** | **15** | **stay in `ROADMAP.md`** — see below | -| **already closed** | **59** | stay as history in the roadmap | - -**Where the 15 ideas live, since "not the register" is not an answer:** they stay in **`ROADMAP.md`**, -which is exactly what that file says it is — *"the prioritized decision log of planned/open work… -Items are intentions"*. The gate exempts them **by their own state word**, so they are not -second-class; they are correctly filed. They are R-6, R-8, R-9, R-12, R-14, R-19, R-34, R-45, R-46, -R-56, R-58, R-62, R-65, R-69, R-72. - ---- - -## 4. The migration — 17 items, and the oldest - -Each moved **verbatim**, nothing added or reinterpreted; the roadmap keeps its copy marked -`MOVED -> OPEN-ITEMS.md` and is not deleted, because that file's job is history. - -| id | originally filed | note | -|---|---|---| -| **R-10** | **2026-07-15** | **the oldest — 38 days** (T-6E-1, dir-fsync asymmetry) | -| R-25, R-40, R-49 | 2026-07-19 | | -| R-30, R-31, R-32, R-35 | 2026-07-21 | three of them marked **[P2-HIGH]** | -| R-78 | 2026-07-25 | an **owed operator decision**, not a defect | -| R-76, R-79 | 2026-07-26 | | -| R-96 | 2026-07-27 | **moved as CLOSED** — both rules are committed at `workspace-CLAUDE.md:48-70` under a heading naming R-96 | -| R-102, R-104, R-105, **R-107** | 2026-07-28 | R-107 **moved as CLOSED** — shipped as R-354 on 2026-08-22 | -| **R-103** | 2026-07-28 | **found by the gate, not by me** — see §5 | - -**Staleness named, not edited away**, as instructed: - -- **R-104** — *partly stale.* The self-heal it calls unreachable **was built**: `resticStep` - escalates to `unlock --remove-all` and retries once (`offbox.go:763-768`), and `unlockStale` runs - before every off-site run and restore (`:1274`). Its premise is also doubtful — the probe is - `restic cat config`, a read that takes no lock. **What remains true:** `ClassifyOffsiteFailure` - (`:179-193`) still has no lock case. -- **R-105** — *partly fixed by its own update.* The `drives` third was traced and populated - 2026-07-28; the other two thirds were not re-verified and are carried as written. - -**No halt.** I checked the three READY findings for urgency before migrating them quietly. R-104 is -the one that reads urgent — *"the tier stays dead until a human runs `restic unlock --remove-all`"* — -and **it is not**, for the reasons above. Nothing in the 31 is both open and urgent. - ---- - -## 5. The gate — and it convicted me before it convicted anything else - -`scripts/one_register_gate.py`, wired into `scripts/repo_gates.py` as **`one-register`** (11 gates now, -all OK). It fails when a `ROADMAP.md` row is **neither an idea nor done** and has **no counterpart row -in `OPEN-ITEMS.md`**. The predicate is the roadmap's **own state column**, so it reads data that -already exists rather than asking anyone to maintain a new marker. - -**The control — plant → convict → remove → pass:** - -| step | result | -|---|---| -| baseline | **PASS** | -| plant an open roadmap-only row | **CONVICTED by name**, rc=1, quoted back | -| remove the plant | **PASS**, file byte-identical (md5) | -| plant an `idea` row | **not convicted** — the exemption is real, not a blanket pass | - -**The control caught my control first.** The first plant did *not* convict, and the counts did not -move at all. Cause: **`ROADMAP.md` has no trailing newline**, so `>>` appended the row onto the last -line and it never started at column 0. A flaw in my test, not the gate. Third session running in which -the instrument caught the operator before the corpus. - -**And the gate then caught three rows my hand-sort missed** — the reason is worth recording because it -is this project's most repeated shape: **my sorting regex matched the whole row, where a finding's -body routinely contains the word "shipped"; the gate matches the state cell only.** - -- **R-103** (`READY — 2026-07-28`) — a genuine finding, mis-read by me as done. **Migrated.** -- **R-23** (`BANKED in full`) and **R-13** (`first slice PROVEN-LIVE`) — this project's own done-words, - which the gate did not know. Vocabulary extended; both are correctly exempt. - -**And on the split it caught a genuine disagreement between the two files:** **R-203** and **R-163** -are recorded closed in the register and still open in the roadmap — R-203 even carries two -contradictory roadmap rows. The register is right in both cases; the roadmap copies are marked -`SUPERSEDED 2026-08-22` with the register's verdict. - -### What the gate cannot see — named, not implied - -1. **A finding filed with the state `idea` escapes.** The state column is a human judgement. -2. **A finding written in prose with no `R-` identifier escapes entirely** — this gate matches ids. - That is the previous session's sweep territory and the template rule *"an enumerated gap becomes a - row"*. -3. **A finding that never reaches the roadmap escapes.** Nothing here reads audits or spikes. -4. It checks a counterpart **exists**, never that the two agree. A stale register row passes. - ---- - -## 6. Housekeeping — sizes before and after - -| file | before | after | | -|---|---|---|---| -| `backlog/OPEN-ITEMS.md` | 691 lines / **672,376 B** | 594 lines / **327,109 B** | **−51%**, and it now holds open work only | -| `backlog/CLOSED-ITEMS.md` | — | 156 lines / 61,580 B | **new sibling**, 128 compressed closed entries | -| `backlog/ROADMAP.md` | 249 lines / **239,306 B** | 160 lines / **78,110 B** | **−67%** | -| `backlog/ROADMAP-HISTORY.md` | — | 116 lines / 28,683 B | **new sibling**, 105 finished items | -| `CONTEXT.md` | 2,580 lines / 217,260 B | 2,607 lines / 219,104 B | **deliberately not compressed** — §7 | - -**A sibling, not the bottom of the file** — appending keeps the byte count and the scroll, which is -the thing being fixed. Closed work compressed to **17%** of its bytes in both files. - -**Nothing was deleted.** Every compressed entry ends `full text: git show fddfe00ce268:`. - -**Load-bearing reasoning was not compressed away.** Rather than judge 134 entries by hand, the -compressor **keeps any sentence stating a rule, a fence or a deliberate refusal, verbatim**, under -**Reasoning kept** — 25 entries carry one. Spot-checked: R-320's kept sentence itself records that its -rule is now standing rule 5 in `workspace-CLAUDE.md`, and R-110's rule is in `felhom.eu/CLAUDE.md` -("The installer publishes by TAG, not by push (R-110)") — **both already homed outside the register, -verified by grep with a negative control.** - -### The compressor moved six rows it should not have — caught and reversed - -**`PARTLY CLOSED` and `OPEN — NOT FIXED` both matched a closed-vocabulary applied to the whole status -field.** R-123, R-190, R-214, R-264, R-295 and **R-352** were moved out of the register. Caught by a -follow-up check in the same session and **restored verbatim from commit `fddfe00ce268`** — not from -the compressed form, because an open row keeps its detail. **The same bug, in the same session, as the -one the gate had.** Filed as **R-378** so the next person to write a status predicate reaches for the -leading verdict rather than the whole field. - ---- - -## 7. Why `CONTEXT.md` was not compressed — a disagreement with the task, stated - -**86% of that file — 187,913 of 217,260 bytes — is one `## Standing rulings` section**, carrying 39 -`S-` ids and 153 bullets under a single heading. **Standing rulings are live reasoning, not finished -work.** - -**This prompt's own §3.4 says the decision log is *"dated, never edited afterwards"*** and is *"the -only place that answers 'has this been proposed before, and why did we say no?'"*. Compressing it -destroys exactly that, and there is no per-ruling delimiter, so a mechanical split risks cutting a -live ruling from its reason — **the failure this whole arc is correcting.** 13 mentions of -`SUPERSEDED` sit inside that blob and cannot be separated from live text safely today. - -**So the problem is navigational, not volumetric, and the fix is structural:** give each ruling a -sub-heading with its `S-` id and date, and it becomes linkable and findable **without a word being -edited**. Filed as **R-377 (LOW)** rather than done, because it is a careful pass of its own. - -The file grew by 1,844 bytes — the new decision-log entry in §8. - ---- - -## 8. Where the hot/bulk decision was recorded — **nowhere. That is the finding.** - -Established by reading, not by citation: it exists as **one unmarked bullet** at -`architecture/01-topology-and-trust.md:150-152`. **No dated entry in the decision log, no `R-` row, no -record of when it was taken.** The only mention in `CONTEXT.md` before today was the entry I wrote -*yesterday about failing to read it*. - -Meanwhile its **consequence** — 40 of 53 templates declare no path — is marked **[FACT]** at -`07-backup-architecture.md:296-299`. **A reader met a marked observation beside an unmarked choice**, -and reasonably asked whether it should be so. Four times. - -**Marker usage, measured** (the prompt said "three times"; the real figure is 45): - -| | before | after | -|---|---|---| -| documents carrying the legend | **1 of 8** (`07`: 10 `[DESIGN]`, 35 `[FACT]`) | **8 of 8** | - -**Done:** - -1. **The legend is carried into all seven** other architecture documents — same wording, no third - marker invented — each stating explicitly that **an unmarked statement means *not yet classified*, - never *observed***. -2. **The hot/bulk split is marked `[DESIGN]`**, with a pointer that is honest about what it can point - at: the decision has **no original date on record**, and the new log entry exists *to give it a - home, not to claim it was decided then*. -3. **A decision-log entry** in `CONTEXT.md` — what was chosen, what follows from it, and **what was - rejected** (moving all-hot data to the data drive), so the same proposal does not return. - -**Deliberately not done, and filed:** the existing statements in those seven documents were **not** -swept into one marker or the other. A wrong mark is worse than none. **R-376 (MEDIUM)** records the -remainder and the rule: mark what a session touches. - ---- - -## 9. The template's closing step, as written - -It had none — `compress` → 0 hits, `CLOSED-ITEMS` → 0, `size before` → 0, with a negative control. -Added as **§N.7**, abridged here: - -> **N.7 Housekeeping — before the report, not after (2026-08-22 ruling)** -> -> 1. **Compress what this session closed.** A closed row keeps its title, the version it shipped in, -> its evidence paths, and any sentence stating a rule. Everything else goes, and it moves to -> `backlog/CLOSED-ITEMS.md`. **Nothing is deleted:** the compressed entry names the commit whose -> `git show` returns the full original text. **Open rows are not touched — their detail is doing a -> job.** -> 2. **Rehome live reasoning before compressing it away.** … **Where it is a decision, mark the -> resulting shape `[DESIGN]` in the architecture document and point it at the log entry.** -> **Losing a reason is how a deliberate design becomes a bug in someone's eyes** — that cost four -> mis-filed defect reports in August 2026 (R-370, R-376). -> 3. **State the register's size in the report, before and after.** A number every session is what -> makes growth visible; prose about tidiness is not a mechanism. - ---- - -## 10. Part 4.2 — the storage default - -**Confirmed correctly filed.** **R-368**, rank **LOW**, open, and it states the corrected position: -the default **does** apply at deploy time (`deploy.html:612` pre-selects `.IsDefault`), the residual -is that it lives in the **template, not the server**, so the API has no default and there is nothing -server-side to test. - -**No row anywhere still asserts the default never applies.** The one occurrence of *"the deploy route -never reads it"* in the register is inside R-368's own quotation of the claim it corrects; the SPEC -carries `[CORRECTED 2026-08-22]` blocks; `grep` across the register, `CLOSED-ITEMS.md` and the SPEC -returns nothing else. - ---- - -## 11. Rows and ceiling - -**Opened:** R-376 (MEDIUM), R-377 (LOW), R-378 (CLOSED same session). -**Migrated in, keeping their original identifiers and dates:** R-10, R-25, R-30, R-31, R-32, R-35, -R-40, R-49, R-76, R-78, R-79, R-96 (closed), R-102, R-103, R-104, R-105, R-107 (closed). -**Restored after a compressor error:** R-123, R-190, R-214, R-264, R-295, R-352. -**Ceiling R-375 → R-378.** - -**CI:** `felhom.eu` **id=384 / run_number=252**, head `ef6ac6fe` — **success**. Commit `ef6ac6f`, -pushed to `main`, all **11** gates OK, tree clean. No other repo touched. - ---- - -## 12. What was dropped, and observations - -**Dropped: nothing from Parts 1–4.** Part 4 (droppable first) and Part 3 both completed. -**One thing deliberately not done and argued rather than obeyed:** compressing `CONTEXT.md` — §7, -filed as R-377. - -**Observations — noticed, not acted on:** - -- **`ROADMAP.md` has no trailing newline.** It broke my own gate control and would break any future - `>>` append. One byte; not fixed here because it belongs to whoever next edits that file - deliberately. -- **R-203 has two contradictory rows in the roadmap** — one open, one shipped. Only the open one was - marked superseded; the duplicate remains as history. -- **The migrated rows are large.** Seventeen verbatim roadmap rows are now in the register, which is - why it fell 51% rather than further. They are open work and keep their detail, by the rule. -- **`OPEN-ITEMS.md` is still 327 KB.** The remaining bulk is open rows with long evidence sections — - that is detail doing a job, and the next reduction comes from closing work, not from editing it. -- **The one-register gate compares existence, not content.** R-203/R-163 were caught only because the - register had *no* row for them after the split; a stale-but-present row would pass. Named in the - gate's own docstring as residual hole 4. +`python3 scripts/repo_gates.py --fast` — all green **except `golden-currency`**, which is correct and +is operator item 1: controller 0.219.0 is released and no golden carries it. diff --git a/STATUS.md b/STATUS.md index c094975f..e156000d 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,7 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-22 — both of last night's worst findings are fixed and proven on the machine.** +**Updated 2026-08-22 — the off-site restore now works for all 53 apps, not 13. It is released and +NOT yet delivered: two steps below are yours.** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates > part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.** @@ -8,211 +9,94 @@ ## Waiting on you -*Nothing. All three questions that stood here were answered on 12–13 August and have moved to -**Decided** below. A decided question left in the deciding list is how a person loses track of what is -actually waiting.* +*This section is allowed to be longer than one screen, and each item says what happens if you do +nothing.* + +1. **Bake and vouch a machine image carrying controller 0.219.0.** Only you can vouch it. + **If you do nothing:** a machine installed today receives 0.218.0 — the release below exists, is + tested and is on the shelf, and no new machine gets it. The build system is red until this is done + and will mail you about it on every push. *(register: R-334's gate, now firing for 0.219.0)* +2. **Then raise the auto-update floor to 0.219.0 — last, in a separate save.** It acts within seconds. + **If you do nothing:** every existing machine stays on 0.218.0, so the fix below reaches nobody and + 40 of 53 apps stay un-restorable on the actual fleet. *(register: R-343's rule)* +3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. + Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as + it is, at the risk you accept by leaving it. I can change it without ever showing you the new one. +4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is + wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, + as it misled one by an hour. ## Decided — and what would reopen each -*A decision with no trigger becomes a permanent silence, so each one names what would make us look -again.* - -- **Getting old backups back yourself: NOT BUILT, deliberately.** A customer in that position is a - support conversation, and we can do it by hand. **Reopens if:** a real customer actually asks — - one request from a person who is not us. *(register: R-312)* -- **The unopenable old copy on `demo-felhom`: KEPT, as a test fixture.** Not for sentiment: it is the - only state in existence where a set-aside store is present and cannot be opened, which is the case - any future handling of lost backups has to face honestly. **Delete it when:** that work ships, or is - abandoned. Until then it is a fixture, not an accumulation. *(register: R-313)* -- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** Its real-world likelihood is - unknown, and hiding one card risks hiding a real second failure. **Reopens if:** it is observed - happening outside a constructed test. *(register: R-303)* +- **Getting old backups back yourself: NOT BUILT, deliberately.** **Reopens if:** a real customer + asks. *(R-312)* +- **The unopenable old copy on `demo-felhom`: KEPT as a test fixture** — the only state in existence + where a set-aside store is present and cannot be opened. **Delete when:** that work ships or is + abandoned. *(R-313)* +- **A machine in two kinds of trouble says both things: LEFT AS IT IS.** **Reopens if:** observed + outside a constructed test. *(R-303)* ## What works -Both demo machines are home, healthy and reporting on the approved pair — **controller 0.217.0, agent -0.130.0**. On `demo-hp` the floor delivered it unaided on 21 August: 0.216.0 → 0.217.0, **17 seconds** -from the operator pressing save to the new controller reporting healthy, with no customer action. -Off-site is credentialed on `demo-hp`, its repository opens with the machine's own key, and -`restic check` over the whole store reports **no errors**. `drill-r50` is reverted to `virgin`, powered -off. **`demo-hp` was reinstalled on 21 August and its off-site backup had been silent since 9 August** -— the rebuild lost the off-site target, and after that was healed every per-app off-site switch was -still off, so the nightly run reported "backup OK" having backed up nothing. Both are on again. +Both demo machines are home, healthy and reporting — agent **0.130.0** published and running on both. +`demo-hp` runs controller **0.219.0**; the fleet floor is still **0.218.0** (see item 2 above). +Off-site is credentialed on `demo-hp` and its store opens with the machine's own key. -**What the fleet actually is, because two summaries have now been misread:** the hub holds **five -customer records and three machines**. The machines are `demo-felhom` and `demo-hp` (both ours, both -disposable) and `drill-r50` (a nested drill VM on DooPlex, reverted and powered off). **`peti-felhom` -is a real machine we have not heard from since 15 July** and has no host record. **`tester-1` is a -record with no machine** — created 13 August, no host, no backups, nothing to lose. +**The fleet, because two summaries have been misread:** five customer records, three machines. +`demo-felhom` and `demo-hp` are ours and disposable; `drill-r50` is a nested drill VM, reverted and +off. **`peti-felhom` is a real machine we have not heard from since 15 July** and has no host record. +**`tester-1` is a record with no machine.** ## Shipped -- **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a - completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks - only ever watched how full a store was getting, and a failed reading was written to a log nobody - reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the - weekly backup that happened to fall inside it. After about half an hour of being unable to see a - store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back. - **Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have - caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That - second check is written down as the next step. - -- **The auto-update floor is current again** (R-343). You raised it to today's version this - afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*, - after the image is vouched, because it acts within seconds. It did: **one machine updated itself - nine seconds later** and came back up cleanly. Worth knowing that raising it is a fleet action, not - paperwork. -- **A dated check can no longer be quietly missed** (R-341). When we write "measure this again on the - 19th", that date is now read by the build system, and a push is refused once it passes. **It is not - a reminder service** — it speaks on the next push, not on the day — and that limit is written into - the check itself. - -- **A new machine installed today finally gets today's software** (R-334, closed). The pre-built image - had been two releases behind since the 14th — anyone installing would have received a version - missing last week's disk-warning fix *and* the follow-up that corrected it. A fresh image was baked - and published, and **you vouched it**, which was the half that could not be done without you. **The - build system is green again for the first time since 14 August**, so the failure mail should stop. - The running machines were not touched: this only ever affected *new* installs. - -- **Last night's two backup alarms were real, and are fixed** (R-336). Both machines failed their - off-site backup at 04:30; **neither machine was at fault**. The off-site box in Germany had run out - of one internal resource and, while looking perfectly healthy from outside, was accepting no - connections at all — for 9½ hours. Restarted, given a ceiling 64× higher, and **both missed backups - were re-run the same morning and are on the off-site box**. **Nothing was lost and nothing was - skipped:** the daily copies on the machines themselves were never affected, and the off-site copy is - weekly, so exactly one attempt fell in the window. The underlying cause — we ask that box a question - about once a second, all day — is filed and **not yet fixed**; the raised ceiling buys **under a - year** on the corrected measurement, not a cure. - -- **The removal now genuinely reverses the installation** (R-316, `installer-v1.28.0` published). The - second reinstall used to hit our own leftover; it was watched failing on the cycle that actually - fails, then watched passing. -- **A correct recovery code is no longer called wrong** (R-311, three components). If a customer types - the code for an older set of backups, the machine now checks the packages we kept, recognises it, and - says so: *your code is correct, it belongs to an earlier package, we kept it, your current backups are - fine, write to us*. It deliberately promises no restore, because there is no button yet. -- **The drive can be re-attached after a reinstall** (R-280), and **the orphan card and the countdown - banner stop promising retrieval they cannot see is still true** (R-294, R-299, R-302). -- **One name per secret — now both halves** (R-295). The dashboard code is „Beállító kód" everywhere, - on the machine *and* in the hub's emails; „Visszaállító kód" is retired. It collided with the escrow - „Helyreállítási kód" and cost a real code. -- **The hub can see whether a machine's guest still has working networking** (R-319, first reader built - against R-264). A machine quietly repairing its own network over and over is now visible instead of - being a green tick; a machine that does not report it is drawn as unknown, never as healthy. -- **The third secret has its own name** (R-323, on your ruling). The five-word phrase that proves an - account owns the box being linked is „Tulajdonosi jelmondat". It was „Visszaállító jelszó" — one word - from the name we retired last week, and false besides: it restores nothing. Five places, all in the - hub; no machine touched. -- **A machine we tell to be quiet is no longer reported as dead** (R-321). It went stale, then down, - then e-mailed you twice about a silence you asked for. It turned out to be two alarms, not one — the - morning backup reminder had the same blind spot and is fixed with it. -- **The hub's own words are under a guard** (R-324). Every customer e-mail and the linking pages are - now checked for a retired name, and the guard has been watched catching one, ignoring an - explanation of one, and going quiet again. -- **The countdown on `demo-felhom` is cancelled** on your ruling (R-307). Nothing was deleted. +- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on + `demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve", + and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they + were never given a drive to choose. It was asking one question to answer two. Proven today on + `privatebin`: data planted through the app itself, backed up, **deleted**, restored — **all 15 files + back byte for byte**, Hungarian accented names included, message „0 fájl és 1 adatkötet + visszaállítva". The 13 apps that do have a drive are unchanged, checked the same way. +- **The off-site restore gives an app's data back at all** (R-354, controller 0.218.0). It used to say + „0 fájl visszaállítva", report success, and the folder was simply not there. It now names what came + back, because a restore that mentions only its file count is how a silent loss reads as a success. +- **Paperless's database is in the backup, and restoring it takes an undo copy first** (R-355, + controller 0.218.0). The dump was landing in a folder named after an app that does not exist. **One + app of 53 was affected**, established with a check first proved able to catch a planted second case. +- **The system tells you when it cannot see the off-site copies** (R-339) — a mail after ~30 minutes, + hourly while it lasts, one all-clear. **Caveat:** it watches whether the machine answers, so it + would *not* have caught the 18 August fault, where one service was wedged and the machine stayed + healthy. +- **The connection leak was ours and is fixed** (R-344). Our agent opened a connection to the off-site + box every 15 minutes and never closed it; the idle timer was switched off. Proved by fixing one + machine and leaving the other: same work, 4 more leaked on the untouched one, none on the fixed one. + The off-site box is back to **17** open connections from **415**. +- **A dated check can no longer be quietly missed** (R-341) — but it speaks on the next push, not on + the day. **A machine we tell to be quiet is no longer reported as dead** (R-321). **One name per + secret** (R-295, R-323). **The hub's own words are under a guard** (R-324). **Removal reverses the + installation** (R-316). **A correct recovery code is no longer called wrong** (R-311). **The drive + can be re-attached after a reinstall** (R-280). ## Broken, or knowingly incomplete -- **FIXED and proven on the machine: the off-site restore gives the app's data back** (R-354, - controller 0.218.0). The same planted files, the same steps, both runs on `demo-hp`: on the old - build the restore said „0 fájl visszaállítva", reported success, and the folder was simply not - there. On the new one it says „**0 fájl és 1 adatkötet visszaállítva**" and all five files come back - **byte for byte**, Hungarian accented names included. The message now names what came back, because - a restore that mentions only its file count is how a silent loss reads as a success. *(register: - R-354, CLOSED)* -- **FIXED and proven on the machine: Paperless's database is in the backup, and a restore of it now - takes an undo copy first** (R-355, controller 0.218.0). The dump was going into a folder named after - an app that does not exist, so nothing collected it — and because the same wrong name was used when - looking for the live database, a restore took **no undo copy at all**. Now: the dump is in the app's - own backup and in the off-site copy for the first time; the restore said „**0 fájl és 3 adatkötet és - az adatbázis visszaállítva**"; and with the undo deliberately made impossible the restore **refused - and did not even stop the app**. We no longer guess which app a database belongs to — Docker already - tells us. **One app of 53 was affected**, established with a check we first proved could catch a - planted second case. *(register: R-355, CLOSED)* -- **STILL BROKEN, and it is now the one that matters most: 40 of our 53 apps still cannot use the - off-site restore at all** (R-356). It refuses before it starts, says a running app „nincs telepítve" - — is not installed — and tells the customer to reinstall it "to the same place", which those apps - give them no way to choose. **Those are exactly the apps whose entire data is the thing R-354 just - fixed**, so today's fix cannot reach them until this one is done. *(register: R-356)* -- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not the - agent. We find out at restore time. A deliberately corrupted copy was detected instantly by the - standard tool — which we never run. *(register: R-359)* - -- **We interrogate the off-site box about once a second** (R-336). Roughly 85,000 questions a day, - for a box we actually write to once a week. That volume is what turned a slow internal leak into - last night's outage in a fortnight. The higher ceiling makes it rare, not impossible — **the leak - itself is untouched.** The honest health check is the resource count climbing, not the absence of an - alarm. **Corrected later the same morning: my first estimate of how fast it leaks was too - optimistic by about 2.5×** — measured properly it is under a year to the new ceiling, not two years. - A deadline, not a comfort. -- **The leak is OURS, not Proxmox's, and asking fewer questions would not have fixed it** (R-344, found - 2026-08-20). We finally looked at who was on the other end of the stuck connections. Every single one - belongs to **our own agent** on the two demo machines — it opens a connection to the off-site box on - each 15-minute cycle and never closes it, and neither does the box. The two Proxmox pollers that make - 99.5% of the traffic leak **nothing at all**. So the plan recorded under R-336 — turn the question - rate down, then watch the count stop climbing — **would have changed nothing and looked like a failed - fix.** The count still says about a year to the ceiling. **Nothing is broken today**; this is a - deadline we now know where to aim at. *(register: R-344, and R-336 re-ranked)* -- **FIXED the same day, and proven on the machines** (R-344). One line of our code: the agent had the - idle-connection timer switched off, so nothing ever retired the connections it abandoned. Turning it - back on to the standard 90 seconds fixed it. **The proof cost nothing clever:** we put the fix on one - machine and left the other alone, and in the same hour the untouched one leaked 4 more connections - while the fixed one leaked none — with both doing exactly the same four rounds of work. **The - off-site box is back to 17 open connections, its normal resting number, down from 415.** All of the - built-up connections released themselves when the agents restarted; the off-site box was only ever - read from, never touched. *(register: R-344)* -- **PUBLISHED the same day, on your word** (R-347, closed). Agent **0.130.0** is released, and the hub - now hands it to any new machine. Both demo machines run the exact published copy. Nothing else on - that screen was changed — in particular the controller floor was left alone, and the "minimum agent" - setting too, because raising that would have **stopped** machines getting updates rather than - helping them. -- **Two things the release itself turned up.** (1) The machines were briefly running a *different* - build of the same version number — harmless here, but nothing in the system would ever have noticed, - because everything compares the version *name*. Now corrected, and filed so it cannot repeat - (R-349). (2) **I printed the hub password into my own session log** while confirming the change - (R-350). It is not in git and not in any saved file — but it is in the log on this machine. - **Changing it is your call**; I can do it without ever showing the new one. Ask and I will. -- **The off-site box was updated, and it did not help — as expected** (R-341). On your ruling we - installed the newer backup software for the practice, having first read its release notes and found - **nothing** about the fault we have. The update went cleanly and everything works, but the leak - behaves exactly as before, which is the result the release notes predicted. **Two dated checks are - booked — 19 August and 25 August** — because half an hour of watching cannot honestly settle it. -- **One machine's status took several minutes to admit a backup had worked** (R-337). `demo-hp` was - still showing this morning's failure for four minutes after the copy was safely on the off-site box; - `demo-felhom` updated in under a minute. **It corrected itself** and both machines now read - correctly, so this is a note to watch, **not something broken** — but during the repair it looked - briefly like a second fault, which is the reason it is written down. -- **`demo-hp` is not set up the way our own notes say it is** (R-338). Our inventory records both - machines as moved onto the isolated internal link last July. `demo-felhom` was; **`demo-hp` was - not** — its control channel is still bound to the ordinary home network, which is exactly what that - change existed to stop. The machine works; the page is wrong, and it misled this session by an hour. - Your call which one to correct. - -- **Peti's machine has no recovery route at all** — see the `PETI` row. **This is a real machine - belonging to a real person**, not one of ours and not a record: it reported to the hub for four and a - half months and has been silent since 15 July, when its host record was deleted. There is no key, no - off-site copy and no local backup. **If that drive fails, everything on it is lost.** First act of the - visit: copy the ~3.6 GB off before anything is reinstalled — it is currently the only copy in - existence. Whether it stays parked is your call and is deliberately left open. -- **Kept backups can be opened — but still only by us** (R-304 partly closed, R-312 decided-not-built). - Today the honest answer is "your code is right, write to us" — and we can. -- **The agent picks dnsmasq by looking at a file another package owns** (R-317). The box installs fine; - only LAN name resolution goes missing, and quietly. One line, deliberately not taken tonight. -- **The storage page has its own separate reason for showing an empty list** (R-298), untouched. -- **Three more facts the machines send still have no reader** (R-264): a staged-but-unapplied agent - update, how deep a restore test actually went, and the two backup-integrity timestamps. Five others - are now recorded as deliberately unread, which is honest rather than fixed. -- **Two thirds of the standing picture is still unproven, and now you can ask** (R-326). - `python3 scripts/unproven.py` lists it: of 55 claims, **23 are walked and 32 are not** — and of - those 32, only 6 point at an evidence document. **The "nine" I have been repeating was wrong**: nine - is how many claims the 9 August review *lowered*, which is a different question. -- **The picture still describes one defect we have since fixed twice** (R-327) — the naming claim. Its - status may only be raised after the capability map moves first, which is a separate judgement. +- **Nothing ever checks that the off-site store is still readable** (R-359). Not the controller, not + the agent. We find out at restore time. A deliberately corrupted copy was caught instantly by the + standard tool — which we never run. +- **A restore that returns nothing still reports success** (part of R-354's neighbourhood, not fixed + today) and **verification copies have no delete guard**. Both deliberately left for their own rows. +- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we + write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling + is **under a year** away on the corrected measurement, not two. +- **Peti's machine has no recovery route at all.** A real machine belonging to a real person, silent + since 15 July, no key, no off-site copy, no local backup. **If that drive fails, everything on it is + lost.** First act of any visit: copy the ~3.6 GB off before anything is reinstalled. +- **The agent picks dnsmasq by looking at a file another package owns** (R-317) — one line; LAN name + resolution goes missing quietly. +- **Three facts the machines send still have no reader** (R-264); **the storage page has its own + reason for an empty list** (R-298); **two thirds of the standing picture is unproven** (R-326: + 23 of 55 claims walked — `python3 scripts/unproven.py`); **the picture still describes one defect we + fixed twice** (R-327). ## Working on next -**One decision is waiting for you:** whether to publish agent 0.130.0 so machines other than the two -demo boxes get the fix (R-347). It needs the artifact screen, which only you can drive. After that: the three remaining -R-264 readers, now that one has been built and we know what one costs; R-317 (one line in the agent); -R-327 (decide what the naming claim's status should be); then the 2026-08-09 batch (R-279 … R-292), -still untriaged against everything since. +The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one +line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store). diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index c2aee2bf..4cb0e2e6 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -334,12 +334,43 @@ refuses **before** stopping the app and names the action that works. |---|---|---|---| | Tier-1 | unit incl. volume tars + DB dumps | all of it | none | | Tier-2 | unit mirror **+** file legs | `hdd/` and `userdata/` **only** (`tier2_restore.go:101-104`) | **the unit mirror is read by nothing** — `RecoveryUnitPath` resolves to `backups/primary/` (`appbackup/paths.go:46-48`) → **R-102** | -| Tier-3 | unit (incl. volume tars) + mandatory legs | files + DB replay; the unit is **skipped** on the way to live (`offbox_reconstitute.go:284-289`; placed only if the live unit is absent, `offbox_restore.go:352-356`) | **no offsite action unpacks the named-volume tars it captures** → **R-107** | +| Tier-3 | unit (incl. volume tars) + mandatory legs | files + DB replay **+ the named-volume tars, replayed from the scratch unit** (`offbox_reconstitute.go` `volReplay`, controller **v0.218.0**); the unit itself is still **skipped** on the way to live (`offbox_reconstitute.go:284-289`; placed only if the live unit is absent, `offbox_restore.go:352-356`) | **CLOSED — R-107** | + +**[FACT] 2026-08-22 — the Tier-3 row above was corrected; the Tier-2 row was NOT.** Until controller +**v0.218.0** this table said *"no offsite action unpacks the named-volume tars it captures"*, and that +was true from the day Tier-3 shipped until 2026-08-21. **R-107 closed in v0.218.0**: `volReplay` +(`offbox_reconstitute.go`) replays the scratch unit's `volume-dumps/` into the live named volumes, +proven live on `demo-hp`. The old sentence is kept here, in the past tense, because a correction that +erases what was believed leaves the next reader no way to tell a fixed gap from one that was never +noticed. + +**R-102 — the Tier-2 half — is NOT closed and nothing in this correction touches it.** The Tier-2 row +above stands exactly as written: the secondary unit mirror is still read by nothing. Do not read +"R-107 closed" as covering both; they were always two register rows, and only one of them moved. **[FACT]** Tier-2's gap is the sharper one because of *when* it bites: Tier-2 exists for the case where the primary drive is lost — and in exactly that case the primary unit is gone while this mirror survives on the second drive, unreachable by any customer action. +**[DESIGN] 2026-08-22 — the restore destination is resolved by the same rule as the capture +destination.** The drive if the app declares one (`HDD_PATH`), the system data path otherwise — +`Manager.GetAppDrivePath`, one expression, used by `CaptureRecoveryUnit` and, since controller +**v0.219.0**, by `ReconstituteFromOffsite` and `PlaceOffsiteRestore` too. + +The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351) applies +to apps that **have a drive to get wrong**. It used to be reached by testing `HDD_PATH == ""`, which +also answered "is this app installed?" — one predicate for two questions. Measured in the catalogue at +`459766cb1639`: **53 templates, 13 declare `needs_hdd: true`, 40 declare `false`**, so for 40 apps that +test was permanently true and the off-site restore refused them forever, while they were running, +with a message telling the customer to reinstall them "in the same place" — a place those apps never +offer. **An app with no drive is not misconfigured** (§8 of `01-topology-and-trust.md` carries the +`[DESIGN]` marker); it is the majority case. + +Since v0.219.0 the two questions are asked separately: *installed?* of `ListDeployedStacks()`, failing +CLOSED when there is no provider to ask; *where?* of `GetAppDrivePath`. A third refusal, with its own +sentence, covers installed-but-no-resolvable-data-root. **R-356**; reasoning also recorded in +`felhom-controller/CONTEXT.md`. + --- ## 7. The recovery chain (D3) — the reason this document exists @@ -525,10 +556,14 @@ anywhere under the backup namespace. Full record: `audits/D5-drive-alone-restore drive. In that failure the primary recovery unit is gone; the surviving mirror on the second drive is `backups/secondary//recovery-unit/`, which **no code path reads** (§6.3). For the 45-or-43 class-B apps the restore is a guaranteed no-op in exactly its designed scenario. → **R-102** -- **Tier-3 vs guest loss.** Tier-3 holds the volume tars and the DB dump. Reconstitution requires - the app to be deployed and skips the unit; the tars are unpacked only by the Tier-1 path, which - requires the guest's secrets. So offsite alone cannot rebuild an app onto a fresh guest. - → **R-107** +- **Tier-3 vs guest loss — HALF of this closed.** Tier-3 holds the volume tars and the DB dump. It + was true until controller v0.218.0 that the tars were unpacked only by the Tier-1 path; since + v0.218.0 the reconstitution replays them itself (`volReplay`) → **R-107 CLOSED 2026-08-22**. What + is **still** true, and is the part that was never R-107: reconstitution requires the app to be + **deployed** and skips the unit, so offsite alone cannot rebuild an app onto a fresh guest. That + requirement is a deliberate decision (R-253) — the restore does not choose a customer's drive for + them — and since **v0.219.0** it means only what it says: an app that is not deployed. It no longer + catches the 40 driveless apps → **R-356**. ### 7.3 ~~What D5 would change — and why it was blocked~~ — **D5 SHIPPED 2026-07-30 (v0.188.0)** @@ -737,7 +772,7 @@ crosses the line — **R-158**. | 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) · **27.6 s** (D5 drill, guest `app.yaml` absent) | 24 h | **PROVEN** (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; **D5 v0.188.0 closed the content gap** — after a restore with the guest's `app.yaml` moved aside, the app read the seeded row **over TCP with its own credential**, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No `.sql` dump in the unit ⇒ the DB came back from the volume tar | | 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery | | 3b | *same, for a class-B app via Tier-2* | — | **no route** — Tier-2 never reads the unit mirror | — | | | **NONE** | §6.3; **R-102** | -| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 or 9 of 53 apps); Tier-3 reconstitute for files + DB; **the volume tars in either copy are unreachable** | **customer** (both) | | 24 h | **PARTIAL** | §7.2; **R-102**, **R-107** | +| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 or 9 of 53 apps); Tier-3 reconstitute for files + DB **+ the named-volume tars since controller v0.218.0** (`volReplay`); **the Tier-2 copy's volume tars remain unreachable** | **customer** (both) | | 24 h | **PARTIAL** | §7.2; **R-102** (open); **R-107 CLOSED 2026-08-22, v0.218.0** | | 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go:359-393` | | 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session | | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | @@ -899,7 +934,8 @@ does **not** hold as written. → **R-108** | **R-104** | An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach, reported as *„ismeretlen okból"* | the offsite tier stays dead until a human unlocks. Was C9-F3 | | **R-105** | Three hub-held DR records are empty on the whole live fleet: `hosts.dr_record_json`, `host_escrow.directive_json`, `dr_recipe.host_half.drives` | the Recipe (§4) is incomplete in exactly the fields host-loss recovery reads. Causes may differ per field | | **R-106** | `dr_recipe.host_half.pbs.namespace` records `"root"` on every box | the recorded restore coordinate is wrong; real namespaces are per-customer | -| **R-107** | No offsite action unpacks the named-volume tars Tier-3 captures on every run | offsite alone cannot rebuild a named-volume app (§7.2) | +| **R-107** | ~~No offsite action unpacks the named-volume tars Tier-3 captures on every run~~ — **CLOSED, controller v0.218.0, 2026-08-22** (`volReplay`, proven live on `demo-hp`). True from the day Tier-3 shipped until 2026-08-21. | was: offsite alone cannot rebuild a named-volume app (§7.2). Now: the tars replay; what remains is that the app must be **deployed** (R-253), which is not R-107 | +| **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) | | ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` | | **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) | | R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10) | diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/README.md b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/README.md new file mode 100644 index 00000000..fe1e00a4 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/README.md @@ -0,0 +1,78 @@ +# DRILL — R-356: the off-site restore for an app with no data drive (2026-08-22) + +**Subject:** `demo-hp` (Tier 0, disposable), guest 9201, controller **v0.219.0**. +**Method:** endpoint-level. `claude-in-chrome` is not available on DooPlex, so every action was the +exact HTTP endpoint the UI's own form posts to — `POST /backup/offbox/run`, +`POST /backup/offbox/restore`, `POST /backup/offbox/reconstitute` — driven with the session cookie and +the session CSRF token scraped from the page. No server logic was skipped; only rendering. + +## What this drill was for + +Nobody had ever driven an off-site reconstitution to completion for an app whose snapshot contains +**only** the recovery unit. The R-356 refusal always fired first. So the downstream legs — placement +mapping, the stat pre-pass, `CheckPlacement`, the safety dump, the DB-service-only window, `volReplay` +— had never run in that shape. **This was a first-run measurement, not a confirmation.** + +**Result: every downstream leg held.** No finding was raised against them. + +## Phase 1 — the driveless class (`privatebin`, `needs_hdd: false`) + +| step | what happened | +|---|---| +| plant | 2 pastes created through PrivateBin's **own JSON API** (`X-Requested-With: JSONHttpRequest`) + 4 files planted directly into the app's named volume, one with a Hungarian accented name | +| findability control | planted a control string, **found** it, removed it, **failed to find** it — `02` in the evidence dir | +| off-site | `POST /backup/offbox/run` → snapshot at 11:16:34Z, 2m20s, `status: ok`, store 44.5 → 47.0 MB | +| destroy | 11 files deleted outright (`rm -rf`, no trash) — `03-destruction.txt` names every one | +| prepare | `POST /backup/offbox/restore mode=full` → snapshot `306accff` into `/mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin` | +| **the scratch shape** | **unit only** — `.../mnt/sys_drive/felhom-data/backups/primary/privatebin/` with `volume-dumps/privatebin_privatebin_data.tar`. The whole dataset is that one tar. This is the shape that had never been driven through. | +| restore | `POST /backup/offbox/reconstitute` — **no refusal**. Volume replayed, app restarted, healthy. | +| verify | **15/15 files byte-for-byte identical**, `diff` of the two sha256 manifests is empty | + +**The accented filename, as explicit hex bytes:** +`c3a17276c3ad7a74c5b172c5912d74c3bc6bc3b67266c3ba72c3b367c3a9702d723335362e747874` +(`árvíztűrő-tükörfúrógép-r356.txt`). Non-ASCII never crossed a shell chain: the planting script was +base64-encoded on DooPlex and decoded inside the guest. + +**The customer-facing message, verbatim (160 bytes):** + +> A(z) privatebin: 0 fájl és 1 adatkötet visszaállítva (mentés: 2026-08-22 13:14) — az alkalmazás +> újraindult. Ennek az alkalmazásnak nincs adatbázisa. + +**Judged, not merely recorded.** It **names the volume** beside the file count, so it is not the R-354 +shape wearing a new hat: a customer reading it is told what actually came back. There is no warning +sitting beside a success. The one honest criticism — the first number a customer reads is `0 fájl`, +and for all 40 driveless apps it will always be 0 — is a wording matter, filed as an observation and +**not acted on here**. + +## Phase 2 — the drive class (`calibre-web`, `needs_hdd: true`, `HDD_PATH=/mnt/felhom-drives/hdd_1`) + +The same walk, to prove the 13-class did not move. The scratch resolved under +`/mnt/felhom-drives/hdd_1/...` — the app's own drive, not the system data path — and the restore +placed there. See `11`–`14` in the evidence dir. + +## Evidence + +`evidence/` holds every artefact, **copied off `demo-hp` at the end of each phase, before any revert**. +Nothing was reverted or torn down during this drill, so no intermediate teardown could have taken it. + +## Deliberately NOT addressed here + +The message that calls an empty restore a success; the delete guard on verification copies; the +absence of any readability check on the off-site store (**R-359**). Separate rows, untouched. + +### Phase 2 result + +**3/3 files back byte for byte**, accented name included (`Örkény-egyperces-r356.txt`, hex +`c396726bc3a96e792d6567797065726365732d723335362e747874`). The restore placed them under +`/mnt/felhom-drives/hdd_1/userdata/media/books/` — the app's own drive. +`/mnt/sys_drive/felhom-data/userdata/media` **does not exist**, so there was no silent misplacement +onto the system disk. + +Message, verbatim (161 bytes): + +> A(z) calibre-web: 3 fájl és 1 adatkötet visszaállítva (mentés: 2026-08-22 13:20) — az alkalmazás +> újraindult. Ennek az alkalmazásnak nincs adatbázisa. + +**No behaviour change for the 13-class was observed.** That is an absence claim, so it does not stand +on the live walk alone: the unit-test red-proof for it plants the system-data fallback into the drive +path and three tests convict it (`TestR356_ScenarioB_*`, `TestR356_ScenarioE_PlaceDriveApp*`). diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/01-planted-manifest.sha256 b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/01-planted-manifest.sha256 new file mode 100644 index 00000000..95c3c298 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/01-planted-manifest.sha256 @@ -0,0 +1,15 @@ +fa28d421a28ed9b6f4b997d5119c2b8adbacff6e2dd119a179b71230f9664934 ./.htaccess +0e9aeb5dd4334bfaacfd51fd5fa5905ea6856b63c375a842bedc1620ce7c16ab ./38/8e/388e3bd405a76985.php +0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 ./DRILL-2026-08-21/SENTINEL.txt +725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a ./DRILL-2026-08-21/binary-1mb.bin +a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 ./DRILL-2026-08-21/nested/őszibarack.md +07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 ./DRILL-2026-08-21/plain.txt +0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea ./DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt +eb8c7a45d319701dc42c13b3175babc17367ce0f6e945386b11d99a4f9ac1d54 ./DRILL-r356-2026-08-22/SENTINEL.txt +4bcf89782290f2057c85ab875b912be47d5d46732f01f09e35221a25401d74f7 ./DRILL-r356-2026-08-22/binary-1mb.bin +576ff7dcc11b2b980bff935b9c35132bca72fbd8155d8e7c1c716d77b2546740 ./DRILL-r356-2026-08-22/nested/plain.txt +2bfa318e42a4514ebd00da28fac8e1ca5f44a674dd72c30e087c168688503835 ./DRILL-r356-2026-08-22/árvíztűrő-tükörfúrógép-r356.txt +f450cdc502162f20c7444576dec29d22140d009845c43e824cda5c3d3b84e320 ./ff/a1/ffa12cc12cd4fa71.php +985ef4b2fca0b01bbbe5bcf25125293926e4d58053a2737ec9c470b0e74c397d ./purge_limiter.php +b5106bc17cb0ea56e7dcf69cc5480def0668670204c24aa8c7ee6db3f8e569eb ./salt.php +9c94379f2de428183db3364c23af55f1dba88c43af95f737ad5926be26e49713 ./traffic_limiter.php diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/02-offsite-runs.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/02-offsite-runs.txt new file mode 100644 index 00000000..ab76b456 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/02-offsite-runs.txt @@ -0,0 +1,18 @@ +Off-site runs during this drill, read from POST /backup/offbox/run + GET /backup/offbox/status. +(The earlier 02-offsite-run-after-plant.json captured empty — the status GET raced the phase change. + Recorded here from the same endpoint plus the controller log, rather than left as a 1-byte file.) + +RUN 1 — after the privatebin plant + last_run : 2026-08-22T11:16:34Z + last_duration: 2m20s + last_error : "" status: ok active: false + repo_size : 44.5 MB -> 47.0 MB + snapshots : 27 + +RUN 2 — after the calibre-web plant + last_run : 2026-08-22T11:22:13Z + last_error : "" status: ok active: false + snapshots : 27 + log excerpt : "backed up privatebin (/mnt/sys_drive/felhom-data/backups/primary/privatebin, 0 mandatory path(s))" + "backed up calibre-web (/mnt/felhom-drives/hdd_1/backups/primary/calibre-web, 1 mandatory path(s))" + ^ the two classes, visible in one pair of lines: system data path vs the app's own drive. diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/03-destruction.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/03-destruction.txt new file mode 100644 index 00000000..78e25a4c --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/03-destruction.txt @@ -0,0 +1,17 @@ +== deleting: +/var/lib/docker/volumes/privatebin_privatebin_data/_data/38/8e/388e3bd405a76985.php +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-2026-08-21/SENTINEL.txt +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-2026-08-21/binary-1mb.bin +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-2026-08-21/nested/őszibarack.md +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-2026-08-21/plain.txt +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-r356-2026-08-22/SENTINEL.txt +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-r356-2026-08-22/binary-1mb.bin +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-r356-2026-08-22/nested/plain.txt +/var/lib/docker/volumes/privatebin_privatebin_data/_data/DRILL-r356-2026-08-22/árvíztűrő-tükörfúrógép-r356.txt +/var/lib/docker/volumes/privatebin_privatebin_data/_data/ff/a1/ffa12cc12cd4fa71.php +== remaining files: +/var/lib/docker/volumes/privatebin_privatebin_data/_data/.htaccess +/var/lib/docker/volumes/privatebin_privatebin_data/_data/purge_limiter.php +/var/lib/docker/volumes/privatebin_privatebin_data/_data/salt.php +/var/lib/docker/volumes/privatebin_privatebin_data/_data/traffic_limiter.php diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/04-prepare-scratch.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/04-prepare-scratch.txt new file mode 100644 index 00000000..cf72a728 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/04-prepare-scratch.txt @@ -0,0 +1,3 @@ +csrf_len=64 +HTTP/2 302 +location: /backups/restore/app?name=privatebin&flash=A+t%C3%A1voli+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl. diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/05-scratch-contents.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/05-scratch-contents.txt new file mode 100644 index 00000000..52f8a1ec --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/05-scratch-contents.txt @@ -0,0 +1,5 @@ +/mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml +/mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml +/mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json +/mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin/mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/06-reconstitute-post.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/06-reconstitute-post.txt new file mode 100644 index 00000000..8e562374 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/06-reconstitute-post.txt @@ -0,0 +1,3 @@ +csrf_len=64 +HTTP/2 302 +location: /backups/restore/app?name=privatebin&flash=A+teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl. diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/07-reconstitute-log.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/07-reconstitute-log.txt new file mode 100644 index 00000000..13111c82 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/07-reconstitute-log.txt @@ -0,0 +1,30 @@ +2026/08/22 11:14:26 backup.go:750: [INFO] [backup] Volume dump: opengist/opengist_opengist_data → 177.0 KB +2026/08/22 11:14:26 backup.go:859: [INFO] [backup] Restarting opengist after volume dump +2026/08/22 11:14:27 backup.go:849: [INFO] [backup] Stopping paperless-ngx for safe volume dump +2026/08/22 11:14:34 backup.go:725: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_postgres_data for paperless-ngx +2026/08/22 11:14:35 backup.go:750: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 68.9 MB +2026/08/22 11:14:35 backup.go:725: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_redis_data for paperless-ngx +2026/08/22 11:14:35 backup.go:750: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 383.0 KB +2026/08/22 11:14:35 backup.go:725: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_data for paperless-ngx +2026/08/22 11:14:36 backup.go:750: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 488.5 KB +2026/08/22 11:14:36 backup.go:859: [INFO] [backup] Restarting paperless-ngx after volume dump +2026/08/22 11:14:47 backup.go:849: [INFO] [backup] Stopping privatebin for safe volume dump +2026/08/22 11:14:47 backup.go:725: [DEBUG] [backup] Dumping volume privatebin_privatebin_data for privatebin +2026/08/22 11:14:47 backup.go:750: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 2.0 MB +2026/08/22 11:14:47 backup.go:859: [INFO] [backup] Restarting privatebin after volume dump +2026/08/22 11:14:48 backup.go:849: [INFO] [backup] Stopping romm for safe volume dump +2026/08/22 11:14:52 backup.go:725: [DEBUG] [backup] Dumping volume romm_romm_config for romm +2026/08/22 11:14:52 backup.go:750: [INFO] [backup] Volume dump: romm/romm_romm_config → 2.5 KB +2026/08/22 11:14:52 backup.go:725: [DEBUG] [backup] Dumping volume romm_romm_db_data for romm +2026/08/22 11:14:53 backup.go:750: [INFO] [backup] Volume dump: romm/romm_romm_db_data → 152.8 MB +2026/08/22 11:14:53 backup.go:725: [DEBUG] [backup] Dumping volume romm_romm_redis_data for romm +2026/08/22 11:14:53 backup.go:750: [INFO] [backup] Volume dump: romm/romm_romm_redis_data → 14.5 MB +2026/08/22 11:14:53 backup.go:859: [INFO] [backup] Restarting romm after volume dump +2026/08/22 11:15:05 backup.go:584: [INFO] [backup] App-data backup completed: 3 databases (422.0 KB total), 6 volume dump(s) (55.518s) +2026/08/22 11:18:22 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/reconstitute +2026/08/22 11:18:22 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/reconstitute from 172.18.0.4:60064 +2026/08/22 11:18:25 restore.go:140: [INFO] [backup] Restoring Docker volume privatebin_privatebin_data for privatebin +2026/08/22 11:18:26 restore.go:169: [DEBUG] [backup] Volume privatebin_privatebin_data restored successfully +2026/08/22 11:18:26 restore.go:174: [INFO] [backup] Restored 1 Docker volume(s) for privatebin +2026/08/22 11:18:34 offbox_reconstitute.go:452: [INFO] [offbox] reconstituted privatebin from snapshot 306accff: 0 file(s) placed, 0 DB dump(s) replayed, safety dump=., skewed=false +2026/08/22 11:18:34 offbox_handlers.go:455: [INFO] [web] off-box reconstitute privatebin completed (async): files=0 dbs=0 snapshot=306accff diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/08-restored-manifest.sha256 b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/08-restored-manifest.sha256 new file mode 100644 index 00000000..95c3c298 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/08-restored-manifest.sha256 @@ -0,0 +1,15 @@ +fa28d421a28ed9b6f4b997d5119c2b8adbacff6e2dd119a179b71230f9664934 ./.htaccess +0e9aeb5dd4334bfaacfd51fd5fa5905ea6856b63c375a842bedc1620ce7c16ab ./38/8e/388e3bd405a76985.php +0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991 ./DRILL-2026-08-21/SENTINEL.txt +725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a ./DRILL-2026-08-21/binary-1mb.bin +a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4 ./DRILL-2026-08-21/nested/őszibarack.md +07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1 ./DRILL-2026-08-21/plain.txt +0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea ./DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt +eb8c7a45d319701dc42c13b3175babc17367ce0f6e945386b11d99a4f9ac1d54 ./DRILL-r356-2026-08-22/SENTINEL.txt +4bcf89782290f2057c85ab875b912be47d5d46732f01f09e35221a25401d74f7 ./DRILL-r356-2026-08-22/binary-1mb.bin +576ff7dcc11b2b980bff935b9c35132bca72fbd8155d8e7c1c716d77b2546740 ./DRILL-r356-2026-08-22/nested/plain.txt +2bfa318e42a4514ebd00da28fac8e1ca5f44a674dd72c30e087c168688503835 ./DRILL-r356-2026-08-22/árvíztűrő-tükörfúrógép-r356.txt +f450cdc502162f20c7444576dec29d22140d009845c43e824cda5c3d3b84e320 ./ff/a1/ffa12cc12cd4fa71.php +985ef4b2fca0b01bbbe5bcf25125293926e4d58053a2737ec9c470b0e74c397d ./purge_limiter.php +b5106bc17cb0ea56e7dcf69cc5480def0668670204c24aa8c7ee6db3f8e569eb ./salt.php +9c94379f2de428183db3364c23af55f1dba88c43af95f737ad5926be26e49713 ./traffic_limiter.php diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/09-outcome-message-verbatim.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/09-outcome-message-verbatim.txt new file mode 100644 index 00000000..81c0974f --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/09-outcome-message-verbatim.txt @@ -0,0 +1,7 @@ +=== CUSTOMER-FACING OUTCOME MESSAGE, VERBATIM === +A(z) privatebin: 0 fájl és 1 adatkötet visszaállítva (mentés: 2026-08-22 13:14) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa. + +=== the same string as UTF-8 hex bytes === +41287a29207072697661746562696e3a20302066c3a16a6c20c3a973203120616461746bc3b674657420766973737a61c3a16c6cc3ad74766120286d656e74c3a9733a20323032362d30382d32322031333a31342920e2809420617a20616c6b616c6d617ac3a17320c3ba6a7261696e64756c742e20456e6e656b20617a20616c6b616c6d617ac3a1736e616b206e696e6373206164617462c3a17a6973612e + +=== byte length: 160 diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/10-controller-log-phase1.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/10-controller-log-phase1.txt new file mode 100644 index 00000000..afcc40a1 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/10-controller-log-phase1.txt @@ -0,0 +1,400 @@ +2026/08/22 11:17:27 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:17:27 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:17:27 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:17:27 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:17:27 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:17:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:17:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:17:32 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:17:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:17:38 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:17:38 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:17:38 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:17:38Z, active_sessions=24 +2026/08/22 11:17:38 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:17:38 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:17:38 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:17:38 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:17:38 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:17:38 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:17:38Z, active_sessions=25 +2026/08/22 11:17:38 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:17:38 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:17:38 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:17:38 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:17:38 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:17:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:17:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:17:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:17:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:17:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:17:42 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:17:43 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (602ms) +2026/08/22 11:17:44 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (623ms) +2026/08/22 11:17:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:17:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 3m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:17:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:17:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:17:52 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:17:54 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:17:54 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:17:54 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:17:54Z, active_sessions=26 +2026/08/22 11:17:54 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:17:54 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:17:54 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:17:54 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:17:54 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:17:54 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/restore +2026/08/22 11:17:54 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/restore from 172.18.0.4:60064 +2026/08/22 11:17:59 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:17:59 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:17:59 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:17:59Z, active_sessions=27 +2026/08/22 11:17:59 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:17:59 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:17:59 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:17:59 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:17:59 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:18:00 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:18:00 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:18:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:18:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:18:02 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 3m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:18:03 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:18:03 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:18:03 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:18:03Z, active_sessions=28 +2026/08/22 11:18:03 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:18:03 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:18:03 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:18:03 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:18:03 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:18:03 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/restore/status +2026/08/22 11:18:03 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/restore/status from 172.18.0.4:60064 +2026/08/22 11:18:03 server.go:636: [WARN] [web] 404 Not Found: GET /backup/restore/status +2026/08/22 11:18:05 offbox_restore.go:270: [INFO] [offbox] restored privatebin (306accff, full=true) → /mnt/felhom-drives/hdd_1/backups/offsite-restore/privatebin +2026/08/22 11:18:05 offbox_handlers.go:372: [INFO] [web] off-box restore privatebin completed (full=true, async) +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true) +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3517MB avail≈22381MB +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26542728 → total=25898MB avail=22381MB used=3517MB (13.6%) +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.7GB used=8.0GB avail=57.2GB (11.6%) +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.5GB avail=884.7GB (0.6%) +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.08 1.32 0.93 3/765 2831" → 1m=1.08 5m=1.32 15m=0.93 +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 52.9°C (hwmon1) +2026/08/22 11:18:11 info.go:13: [DEBUG] [system] GetInfo done in 59ms — mem=3517MB/25898MB (13.6%), rootDisk=8.0GB/68.7GB (11.6%), load=1.08/1.32/0.93, temp=52.9°C (hwmon1), cpu=10.2% +2026/08/22 11:18:12 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan +2026/08/22 11:18:12 scheduler.go:67: [DEBUG] [scheduler] job stack-scan: execution starting +2026/08/22 11:18:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:18:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:18:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:18:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=true composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "cloudflared" deployed=false composePath=/opt/docker/stacks/cloudflared/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=true composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/08/22 11:18:12 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health +2026/08/22 11:18:12 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=true composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=true composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=true composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=false composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/08/22 11:18:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/08/22 11:18:12 manager.go:1497: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/08/22 11:18:12 manager.go:539: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/08/22 11:18:12 manager.go:570: [INFO] [stacks] ScanStacks complete: 56 stacks found (6 deployed, 50 available) +2026/08/22 11:18:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:12 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 47ms) +2026/08/22 11:18:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 3m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:18:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:12 client.go:1142: [DEBUG] [agentapi] GET /storage -> 200 (577ms) +2026/08/22 11:18:12 scheduler.go:363: [INFO] [scheduler] Job agent-channel-health completed (took 577ms) +2026/08/22 11:18:13 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (607ms) +2026/08/22 11:18:14 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (676ms) +2026/08/22 11:18:21 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:18:21 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:18:21 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:18:21Z, active_sessions=29 +2026/08/22 11:18:21 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:18:21 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:18:21 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:18:21 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:18:21 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:18:21 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:18:21 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:18:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:18:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:18:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 3m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 3m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:18:22 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:22 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:18:22 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:18:22 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:18:22Z, active_sessions=30 +2026/08/22 11:18:22 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:18:22 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:18:22 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:18:22 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:18:22 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:18:22 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/reconstitute +2026/08/22 11:18:22 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/reconstitute from 172.18.0.4:60064 +2026/08/22 11:18:25 dbdump.go:92: [DEBUG] DiscoverDatabases: running docker ps to find database containers +2026/08/22 11:18:25 dbdump.go:111: [DEBUG] DiscoverDatabases: docker ps output: 4c72492af5ca romm romm rommapp/romm:5.0.0 +b261d717618d romm-db romm mariadb:11.4 +9010a9d7927f romm-redis romm redis:7-alpine +83d6e8adc484 privatebin privatebin privatebin/pdo:2.0.5 +80fc6005741e paperless-webserver paperless-ngx ghcr.io/paperless-ngx/paperless-ngx:2.20.15 +ae015380d719 paperless-postgres paperless-ngx postgres:16-alpine +fddae44de81e paperless-redis paperless-ngx redis:7-alpine +3aae6ced0216 opengist opengist ghcr.io/thomiceli/opengist:1.13 +d0c84add6adf kimai kimai kimai/kimai2:apac... +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container romm (image=rommapp/romm:5.0.0, not a database) +2026/08/22 11:18:25 dbdump.go:140: [DEBUG] DiscoverDatabases: found mariadb container: romm-db (id=b261d717618d) +2026/08/22 11:18:25 dbdump.go:160: [DEBUG] DiscoverDatabases: romm-db → stack=romm, dbUser=root, dbName=romm +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container romm-redis (image=redis:7-alpine, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container privatebin (image=privatebin/pdo:2.0.5, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database) +2026/08/22 11:18:25 dbdump.go:140: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=ae015380d719) +2026/08/22 11:18:25 dbdump.go:805: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless") +2026/08/22 11:18:25 dbdump.go:160: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container opengist (image=ghcr.io/thomiceli/opengist:1.13, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container kimai (image=kimai/kimai2:apache-2.57.0, not a database) +2026/08/22 11:18:25 dbdump.go:140: [DEBUG] DiscoverDatabases: found mariadb container: kimai-db (id=fb9ebd95951c) +2026/08/22 11:18:25 dbdump.go:160: [DEBUG] DiscoverDatabases: kimai-db → stack=kimai, dbUser=root, dbName=kimai +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container calibre-web (image=crocodilestick/calibre-web-automated:v4.0.6, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.219.0, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container cloudflared (image=cloudflare/cloudflared:2026.6.0, not a database) +2026/08/22 11:18:25 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/08/22 11:18:25 dbdump.go:167: [DEBUG] DiscoverDatabases: found 3 database(s), skipped 12 non-DB container(s) +2026/08/22 11:18:25 dbdump.go:170: [INFO] [backup] Discovered 3 databases +2026/08/22 11:18:25 manager.go:1067: [DEBUG] [stacks] StopStack privatebin: current state=running deployed=true containers=1 +2026/08/22 11:18:25 manager.go:1070: [INFO] [stacks] Stopping stack: privatebin +2026/08/22 11:18:25 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, SUBDOMAIN, DOMAIN, IMPORT_PATH] (8 app + 0 system) +2026/08/22 11:18:25 manager.go:1270: [DEBUG] Running: docker compose down (in /opt/docker/stacks/privatebin) +2026/08/22 11:18:25 manager.go:1290: [DEBUG] Command completed: docker compose down (took 0.3s) +2026/08/22 11:18:25 manager.go:1079: [INFO] [stacks] Stack privatebin stopped successfully (took 0.3s) +2026/08/22 11:18:25 manager.go:621: [INFO] [stacks] Status refresh: 13 containers across 56 stacks +2026/08/22 11:18:25 restore.go:140: [INFO] [backup] Restoring Docker volume privatebin_privatebin_data for privatebin +2026/08/22 11:18:26 restore.go:169: [DEBUG] [backup] Volume privatebin_privatebin_data restored successfully +2026/08/22 11:18:26 restore.go:174: [INFO] [backup] Restored 1 Docker volume(s) for privatebin +2026/08/22 11:18:26 manager.go:986: [DEBUG] [stacks] StartStack privatebin: current state=stopped deployed=true +2026/08/22 11:18:26 manager.go:989: [INFO] [stacks] Starting stack: privatebin +2026/08/22 11:18:26 manager.go:996: [DEBUG] [stacks] StartStack privatebin: prepared 8 env vars for compose +2026/08/22 11:18:26 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, SUBDOMAIN, IMPORT_PATH] (8 app + 0 system) +2026/08/22 11:18:26 manager.go:1270: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/privatebin) +2026/08/22 11:18:26 manager.go:1290: [DEBUG] Command completed: docker compose up -d (took 0.3s) +2026/08/22 11:18:26 manager.go:1004: [INFO] [stacks] Stack privatebin started successfully (took 0.3s) +2026/08/22 11:18:26 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:29 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, SUBDOMAIN, IMPORT_PATH] (8 app + 0 system) +2026/08/22 11:18:29 manager.go:1270: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/privatebin) +2026/08/22 11:18:29 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:29 restore.go:204: [DEBUG] [backup] Post-restore health check: privatebin not yet running, waiting... +2026/08/22 11:18:29 manager.go:1290: [DEBUG] Command completed: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (took 0.1s) +2026/08/22 11:18:29 manager.go:1350: [INFO] [stacks] Stack privatebin post-start status: +2026/08/22 11:18:29 manager.go:1353: [INFO] [stacks] privatebin privatebin/pdo:2.0.5 running Up 3 seconds (health: starting) +2026/08/22 11:18:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:18:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 4m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (4 skipped not due, 1 skipped no container) +2026/08/22 11:18:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:18:32 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:34 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:34 restore.go:199: [DEBUG] [backup] Post-restore health check: privatebin is running +2026/08/22 11:18:34 offbox_reconstitute.go:452: [INFO] [offbox] reconstituted privatebin from snapshot 306accff: 0 file(s) placed, 0 DB dump(s) replayed, safety dump=., skewed=false +2026/08/22 11:18:34 offbox_handlers.go:455: [INFO] [web] off-box reconstitute privatebin completed (async): files=0 dbs=0 snapshot=306accff +2026/08/22 11:18:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:18:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:18:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:18:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 4m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 4m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 1 targets (4 skipped not due, 1 skipped no container) +2026/08/22 11:18:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:18:42 healthprobe.go:153: [DEBUG] Health probe privatebin: HTTP GET :8080/ → 200 (5ms) +2026/08/22 11:18:42 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:42 healthprobe.go:133: [INFO] Health probes: 1 ok (of 1 probed) +2026/08/22 11:18:42 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:18:42 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:18:42 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:18:42Z, active_sessions=31 +2026/08/22 11:18:42 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:18:42 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:18:42 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:18:42 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:18:42 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:18:42 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:18:42 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:18:43 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (589ms) +2026/08/22 11:18:44 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (601ms) +2026/08/22 11:18:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:18:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:18:52 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:18:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 4m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 4m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:18:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:19:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:19:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 4m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 4m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 4m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 3m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:19:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:19:02 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:19:04 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:19:04 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:19:04 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:19:04Z, active_sessions=32 +2026/08/22 11:19:04 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:19:04 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:19:04 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:19:04 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:19:04 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:19:04 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:19:04 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:19:06 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:19:06 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:19:06 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:19:06Z, active_sessions=33 +2026/08/22 11:19:06 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:19:06 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:19:06 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:19:06 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:19:06 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:19:06 auth.go:134: [DEBUG] [web] auth: valid session for GET /backups/restore/app +2026/08/22 11:19:06 server.go:393: [DEBUG] [web] ServeHTTP: GET /backups/restore/app from 172.18.0.4:60064 +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true) +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3600MB avail≈22298MB +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26444532 → total=25898MB avail=22298MB used=3600MB (13.9%) +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.7GB used=8.0GB avail=57.2GB (11.6%) +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.5GB avail=884.7GB (0.6%) +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readLoadAvg: raw="0.95 1.25 0.92 6/825 3250" → 1m=0.95 5m=1.25 15m=0.92 +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 57.8°C (hwmon2) +2026/08/22 11:19:11 info.go:13: [DEBUG] [system] GetInfo done in 59ms — mem=3600MB/25898MB (13.9%), rootDisk=8.0GB/68.7GB (11.6%), load=0.95/1.25/0.92, temp=57.8°C (hwmon2), cpu=8.4% +2026/08/22 11:19:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:19:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:19:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:19:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:19:12 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health +2026/08/22 11:19:12 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting +2026/08/22 11:19:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:19:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 4m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 4m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 4m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 3m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:19:12 client.go:1142: [DEBUG] [agentapi] GET /storage -> 200 (561ms) +2026/08/22 11:19:12 scheduler.go:363: [INFO] [scheduler] Job agent-channel-health completed (took 561ms) +2026/08/22 11:19:13 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (660ms) +2026/08/22 11:19:14 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (674ms) +2026/08/22 11:19:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:19:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 4m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 3m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 4m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 4m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:19:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:19:22 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:19:25 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:19:25 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:19:25 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:19:25Z, active_sessions=34 +2026/08/22 11:19:25 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:19:25 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:19:25 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:19:25 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:19:25 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:19:25 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:19:25 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:19:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:19:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 5m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 4m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 4m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 3m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:19:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:19:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:19:32 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/11-13class-planted.sha256 b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/11-13class-planted.sha256 new file mode 100644 index 00000000..8d40db0b --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/11-13class-planted.sha256 @@ -0,0 +1,3 @@ +bc6d35efa862246b543a1e66b2a2ab4fa889bbe28a827cc30018593ca026561a /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/book.bin +93defa514998c95ec4e1483b7b4935166ab8f9c6d850b464c39114118360f82d /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/Örkény-egyperces-r356.txt +85c29e22129c14d03815e107fc1fceab004bf682dc030d927f24fdd68fc6ceea /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/SENTINEL.txt diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/12-13class-destruction.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/12-13class-destruction.txt new file mode 100644 index 00000000..8b7f0d91 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/12-13class-destruction.txt @@ -0,0 +1,13 @@ +/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/SENTINEL.txt +/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/book.bin +/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/Örkény-egyperces-r356.txt +-- after: +total 456 +drwxrwsr-x 4 1000 1000 4096 Aug 22 11:22 . +drwxrwsr-x 10 root 1000 4096 Jul 26 06:17 .. +drwxrwxr-x 3 1000 1000 4096 Aug 21 20:09 DRILL-2026-08-21 +-rw-r--r-- 1 1000 1000 181 Aug 4 12:53 DRILL-SENTINEL.txt +-rw-r--r-- 1 1000 1000 413696 Aug 4 12:52 metadata.db +-rw-r--r-- 1 1000 1000 32768 Aug 22 11:20 metadata.db-shm +-rw-r--r-- 1 1000 1000 0 Aug 9 02:15 metadata.db-wal +drwxrwxr-x 3 1000 1000 4096 Aug 9 07:59 rehearsal-2026-08-09 diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/13-13class-scratch.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/13-13class-scratch.txt new file mode 100644 index 00000000..235efe16 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/13-13class-scratch.txt @@ -0,0 +1,20 @@ +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/.felhom.yml +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/app.yaml +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/docker-compose.yml +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/volume-dumps/calibre-web_calibre_web_config.tar +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/POST-SNAPSHOT.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/SENTINEL.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/binary-1mb.bin +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/nested/őszibarack.md +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/plain.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-SENTINEL.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/SENTINEL.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/book.bin +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/Örkény-egyperces-r356.txt +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-shm +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-wal +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/binary-3mb.bin +/mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web/mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/nested/őszibarack.md diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/14-13class-restored.sha256 b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/14-13class-restored.sha256 new file mode 100644 index 00000000..d483f62f --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/14-13class-restored.sha256 @@ -0,0 +1,3 @@ +85c29e22129c14d03815e107fc1fceab004bf682dc030d927f24fdd68fc6ceea ./SENTINEL.txt +bc6d35efa862246b543a1e66b2a2ab4fa889bbe28a827cc30018593ca026561a ./book.bin +93defa514998c95ec4e1483b7b4935166ab8f9c6d850b464c39114118360f82d ./Örkény-egyperces-r356.txt diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/15-13class-outcome-message.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/15-13class-outcome-message.txt new file mode 100644 index 00000000..4b15ce60 --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/15-13class-outcome-message.txt @@ -0,0 +1,7 @@ +=== 13-CLASS OUTCOME MESSAGE, VERBATIM === +A(z) calibre-web: 3 fájl és 1 adatkötet visszaállítva (mentés: 2026-08-22 13:20) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa. + +=== UTF-8 hex === +41287a292063616c696272652d7765623a20332066c3a16a6c20c3a973203120616461746bc3b674657420766973737a61c3a16c6cc3ad74766120286d656e74c3a9733a20323032362d30382d32322031333a32302920e2809420617a20616c6b616c6d617ac3a17320c3ba6a7261696e64756c742e20456e6e656b20617a20616c6b616c6d617ac3a1736e616b206e696e6373206164617462c3a17a6973612e + +=== byte length: 161 diff --git a/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/16-controller-log-phase2.txt b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/16-controller-log-phase2.txt new file mode 100644 index 00000000..b10eeefd --- /dev/null +++ b/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/16-controller-log-phase2.txt @@ -0,0 +1,500 @@ +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=true composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/08/22 11:22:12 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health +2026/08/22 11:22:12 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=true composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=true composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=true composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=false composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/08/22 11:22:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/08/22 11:22:12 manager.go:1497: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/08/22 11:22:12 manager.go:539: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/08/22 11:22:12 manager.go:570: [INFO] [stacks] ScanStacks complete: 56 stacks found (6 deployed, 50 available) +2026/08/22 11:22:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:22:12 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 44ms) +2026/08/22 11:22:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:22:12 client.go:1142: [DEBUG] [agentapi] GET /storage -> 200 (557ms) +2026/08/22 11:22:12 scheduler.go:363: [INFO] [scheduler] Job agent-channel-health completed (took 557ms) +2026/08/22 11:22:13 settings.go:697: [DEBUG] [settings] saved to /opt/docker/felhom-controller/data/settings.json (5736 bytes) +2026/08/22 11:22:13 settings.go:700: [INFO] [settings] Settings saved +2026/08/22 11:22:13 settings.go:697: [DEBUG] [settings] saved to /opt/docker/felhom-controller/data/settings.json (5730 bytes) +2026/08/22 11:22:13 settings.go:700: [INFO] [settings] Settings saved +2026/08/22 11:22:13 offbox.go:1129: [INFO] [offbox] backup OK: 6 app(s) backed up, 27 snapshot(s), 2m8s +2026/08/22 11:22:13 offbox_progress.go:342: [INFO] [offbox] manual run progress reporting ended after 2m14s +2026/08/22 11:22:13 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (679ms) +2026/08/22 11:22:14 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:36778 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:14 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:14 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:14 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:14Z, active_sessions=54 +2026/08/22 11:22:14 auth.go:214: [INFO] [web] Login from 172.18.0.4:36778 +2026/08/22 11:22:14 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:14 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:36778 +2026/08/22 11:22:14 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:14 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:36778 +2026/08/22 11:22:14 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:14 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:14Z, active_sessions=55 +2026/08/22 11:22:14 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:14 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:14 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:14 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:14 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:14 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:14 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:14 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:14 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:14 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (681ms) +2026/08/22 11:22:14 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:15 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:15 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:15Z, active_sessions=56 +2026/08/22 11:22:15 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:15 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:15 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:15 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:15 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:15 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:15Z, active_sessions=57 +2026/08/22 11:22:15 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:15 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:15 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:15 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:15 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:22:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 1m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:22:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 1m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:22:22 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:22:28 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:28 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:28 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:28Z, active_sessions=58 +2026/08/22 11:22:28 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:28 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:28 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:28 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:28 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:28 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/restore +2026/08/22 11:22:28 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/restore from 172.18.0.4:60064 +2026/08/22 11:22:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:22:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:22:32 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:22:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 1m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 1m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:22:35 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:35 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:35 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:35Z, active_sessions=59 +2026/08/22 11:22:35 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:35 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:35 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:35 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:35 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:35 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:35 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:36 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:36 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:36 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:36Z, active_sessions=60 +2026/08/22 11:22:36 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:36 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:36 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:36 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:36 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:36 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:36 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:39 offbox_restore.go:270: [INFO] [offbox] restored calibre-web (b921bae3, full=true) → /mnt/felhom-drives/hdd_1/backups/offsite-restore/calibre-web +2026/08/22 11:22:39 offbox_handlers.go:372: [INFO] [web] off-box restore calibre-web completed (full=true, async) +2026/08/22 11:22:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:22:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:22:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:22:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:22:42 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:22:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 1m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 1m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:22:43 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (605ms) +2026/08/22 11:22:44 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (596ms) +2026/08/22 11:22:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:22:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:22:52 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:22:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:22:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:22:56 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:56 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:56 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:56Z, active_sessions=61 +2026/08/22 11:22:56 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:56 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:56 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:56 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:56 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:56 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:56 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:22:57 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:22:57 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:22:57 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:22:57Z, active_sessions=62 +2026/08/22 11:22:57 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:22:57 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:22:57 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:22:57 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:22:57 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:22:57 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:22:57 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:23:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:23:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:23:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:23:02 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:02 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:23:02 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:23:02 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:23:02Z, active_sessions=63 +2026/08/22 11:23:02 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:23:02 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:23:02 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:23:02 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:23:02 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:23:02 auth.go:134: [DEBUG] [web] auth: valid session for POST /backup/offbox/reconstitute +2026/08/22 11:23:02 server.go:393: [DEBUG] [web] ServeHTTP: POST /backup/offbox/reconstitute from 172.18.0.4:60064 +2026/08/22 11:23:04 dbdump.go:92: [DEBUG] DiscoverDatabases: running docker ps to find database containers +2026/08/22 11:23:04 dbdump.go:111: [DEBUG] DiscoverDatabases: docker ps output: 5afaefbd4e91 romm romm rommapp/romm:5.0.0 +79cae59f7986 romm-db romm mariadb:11.4 +c7c079ae23af romm-redis romm redis:7-alpine +adca811fb993 privatebin privatebin privatebin/pdo:2.0.5 +9fe7dc9f36ae paperless-webserver paperless-ngx ghcr.io/paperless-ngx/paperless-ngx:2.20.15 +4a432b653108 paperless-postgres paperless-ngx postgres:16-alpine +950113e71f40 paperless-redis paperless-ngx redis:7-alpine +4d81ef1c6cde opengist opengist ghcr.io/thomiceli/opengist:1.13 +4df3efdb8366 kimai kimai kimai/kimai2:apac... +2026/08/22 11:23:04 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container romm (image=rommapp/romm:5.0.0, not a database) +2026/08/22 11:23:04 dbdump.go:140: [DEBUG] DiscoverDatabases: found mariadb container: romm-db (id=79cae59f7986) +2026/08/22 11:23:04 dbdump.go:160: [DEBUG] DiscoverDatabases: romm-db → stack=romm, dbUser=root, dbName=romm +2026/08/22 11:23:04 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container romm-redis (image=redis:7-alpine, not a database) +2026/08/22 11:23:04 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container privatebin (image=privatebin/pdo:2.0.5, not a database) +2026/08/22 11:23:04 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database) +2026/08/22 11:23:04 dbdump.go:140: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=4a432b653108) +2026/08/22 11:23:04 dbdump.go:805: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless") +2026/08/22 11:23:05 dbdump.go:160: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database) +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container opengist (image=ghcr.io/thomiceli/opengist:1.13, not a database) +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container kimai (image=kimai/kimai2:apache-2.57.0, not a database) +2026/08/22 11:23:05 dbdump.go:140: [DEBUG] DiscoverDatabases: found mariadb container: kimai-db (id=d3e87cbf6539) +2026/08/22 11:23:05 dbdump.go:160: [DEBUG] DiscoverDatabases: kimai-db → stack=kimai, dbUser=root, dbName=kimai +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container calibre-web (image=crocodilestick/calibre-web-automated:v4.0.6, not a database) +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.219.0, not a database) +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container cloudflared (image=cloudflare/cloudflared:2026.6.0, not a database) +2026/08/22 11:23:05 dbdump.go:133: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/08/22 11:23:05 dbdump.go:167: [DEBUG] DiscoverDatabases: found 3 database(s), skipped 12 non-DB container(s) +2026/08/22 11:23:05 dbdump.go:170: [INFO] [backup] Discovered 3 databases +2026/08/22 11:23:05 manager.go:1067: [DEBUG] [stacks] StopStack calibre-web: current state=running deployed=true containers=1 +2026/08/22 11:23:05 manager.go:1070: [INFO] [stacks] Stopping stack: calibre-web +2026/08/22 11:23:05 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, HDD_PATH, SUBDOMAIN, DOMAIN, USERDATA_PATH, IMPORT_PATH] (10 app + 0 system) +2026/08/22 11:23:05 manager.go:1270: [DEBUG] Running: docker compose down (in /opt/docker/stacks/calibre-web) +2026/08/22 11:23:09 manager.go:1290: [DEBUG] Command completed: docker compose down (took 4.0s) +2026/08/22 11:23:09 manager.go:1079: [INFO] [stacks] Stack calibre-web stopped successfully (took 4.0s) +2026/08/22 11:23:09 manager.go:621: [INFO] [stacks] Status refresh: 13 containers across 56 stacks +2026/08/22 11:23:09 restore.go:140: [INFO] [backup] Restoring Docker volume calibre-web_calibre_web_config for calibre-web +2026/08/22 11:23:09 restore.go:169: [DEBUG] [backup] Volume calibre-web_calibre_web_config restored successfully +2026/08/22 11:23:09 restore.go:174: [INFO] [backup] Restored 1 Docker volume(s) for calibre-web +2026/08/22 11:23:09 manager.go:986: [DEBUG] [stacks] StartStack calibre-web: current state=stopped deployed=true +2026/08/22 11:23:09 manager.go:989: [INFO] [stacks] Starting stack: calibre-web +2026/08/22 11:23:09 manager.go:996: [DEBUG] [stacks] StartStack calibre-web: prepared 10 env vars for compose +2026/08/22 11:23:09 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, HDD_PATH, SUBDOMAIN, USERDATA_PATH, IMPORT_PATH] (10 app + 0 system) +2026/08/22 11:23:09 manager.go:1270: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/calibre-web) +2026/08/22 11:23:09 manager.go:1290: [DEBUG] Command completed: docker compose up -d (took 0.3s) +2026/08/22 11:23:09 manager.go:1004: [INFO] [stacks] Stack calibre-web started successfully (took 0.3s) +2026/08/22 11:23:09 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true) +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3326MB avail≈22572MB +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26768972 → total=25898MB avail=22572MB used=3326MB (12.8%) +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.7GB used=8.0GB avail=57.3GB (11.6%) +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.5GB avail=884.7GB (0.6%) +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.11 1.44 1.09 3/776 5238" → 1m=1.11 5m=1.44 15m=1.09 +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 55.9°C (hwmon2) +2026/08/22 11:23:11 info.go:13: [DEBUG] [system] GetInfo done in 62ms — mem=3326MB/25898MB (12.8%), rootDisk=8.0GB/68.7GB (11.6%), load=1.11/1.44/1.09, temp=55.9°C (hwmon2), cpu=5.8% +2026/08/22 11:23:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:23:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:23:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:23:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:23:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (4 skipped not due, 1 skipped no container) +2026/08/22 11:23:12 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health +2026/08/22 11:23:12 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting +2026/08/22 11:23:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:12 client.go:1142: [DEBUG] [agentapi] GET /storage -> 200 (580ms) +2026/08/22 11:23:12 scheduler.go:363: [INFO] [scheduler] Job agent-channel-health completed (took 580ms) +2026/08/22 11:23:12 manager.go:1261: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, HDD_PATH, SUBDOMAIN, USERDATA_PATH, IMPORT_PATH] (10 app + 0 system) +2026/08/22 11:23:12 manager.go:1270: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/calibre-web) +2026/08/22 11:23:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:12 restore.go:204: [DEBUG] [backup] Post-restore health check: calibre-web not yet running, waiting... +2026/08/22 11:23:12 manager.go:1290: [DEBUG] Command completed: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (took 0.1s) +2026/08/22 11:23:12 manager.go:1350: [INFO] [stacks] Stack calibre-web post-start status: +2026/08/22 11:23:12 manager.go:1353: [INFO] [stacks] calibre-web crocodilestick/calibre-web-automated:v4.0.6 running Up 3 seconds (health: starting) +2026/08/22 11:23:14 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (713ms) +2026/08/22 11:23:14 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (637ms) +2026/08/22 11:23:17 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:17 restore.go:204: [DEBUG] [backup] Post-restore health check: calibre-web not yet running, waiting... +2026/08/22 11:23:18 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:23:18 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:23:18 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:23:18Z, active_sessions=64 +2026/08/22 11:23:18 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:23:18 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:23:18 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:23:18 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:23:18 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:23:18 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:23:18 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:23:18 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:23:18 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:23:18 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:23:18Z, active_sessions=65 +2026/08/22 11:23:18 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:23:19 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:23:19 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:23:19 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:23:19 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:23:19 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:23:19 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:23:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:23:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:23:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (4 skipped not due, 1 skipped no container) +2026/08/22 11:23:22 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:22 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:22 restore.go:204: [DEBUG] [backup] Post-restore health check: calibre-web not yet running, waiting... +2026/08/22 11:23:27 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:27 restore.go:199: [DEBUG] [backup] Post-restore health check: calibre-web is running +2026/08/22 11:23:27 offbox_reconstitute.go:452: [INFO] [offbox] reconstituted calibre-web from snapshot b921bae3: 3 file(s) placed, 0 DB dump(s) replayed, safety dump=., skewed=false +2026/08/22 11:23:27 offbox_handlers.go:455: [INFO] [web] off-box reconstitute calibre-web completed (async): files=3 dbs=0 snapshot=b921bae3 +2026/08/22 11:23:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:23:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:23:32 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 1 targets (4 skipped not due, 1 skipped no container) +2026/08/22 11:23:32 healthprobe.go:153: [DEBUG] Health probe calibre-web: HTTP GET :8083/ → 302 (4ms) +2026/08/22 11:23:32 healthprobe.go:133: [INFO] Health probes: 1 ok (of 1 probed) +2026/08/22 11:23:39 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:23:39 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:23:39 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:23:39Z, active_sessions=66 +2026/08/22 11:23:39 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:23:39 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:23:39 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:23:39 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:23:39 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:23:39 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:23:39 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:23:40 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:23:40 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:23:40 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:23:40Z, active_sessions=67 +2026/08/22 11:23:40 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:23:40 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:23:40 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:23:40 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:23:40 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:23:40 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:23:40 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:23:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:23:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:23:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:23:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:23:42 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:23:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:23:43 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (622ms) +2026/08/22 11:23:44 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (609ms) +2026/08/22 11:23:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:23:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:23:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m20s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:23:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:23:52 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:24:00 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:24:00 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:24:00 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:24:00Z, active_sessions=68 +2026/08/22 11:24:00 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:24:00 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:24:00 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:24:00 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:24:00 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:24:00 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:24:00 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:24:01 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:24:01 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:24:01 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:24:01Z, active_sessions=69 +2026/08/22 11:24:01 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:24:01 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:24:01 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:24:01 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:24:01 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:24:01 auth.go:134: [DEBUG] [web] auth: valid session for GET /backup/offbox/status +2026/08/22 11:24:01 server.go:393: [DEBUG] [web] ServeHTTP: GET /backup/offbox/status from 172.18.0.4:60064 +2026/08/22 11:24:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:24:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:24:02 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:24:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 3m0s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 30s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:24:09 auth.go:146: [DEBUG] [web] login attempt from 172.18.0.4:60064 (X-Forwarded-For: 172.18.0.1) +2026/08/22 11:24:09 auth.go:194: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/08/22 11:24:09 auth.go:260: [DEBUG] [web] session created, expires=2026-08-29T11:24:09Z, active_sessions=70 +2026/08/22 11:24:09 auth.go:214: [INFO] [web] Login from 172.18.0.4:60064 +2026/08/22 11:24:09 auth.go:134: [DEBUG] [web] auth: valid session for GET / +2026/08/22 11:24:09 server.go:393: [DEBUG] [web] ServeHTTP: GET / from 172.18.0.4:60064 +2026/08/22 11:24:09 auth.go:134: [DEBUG] [web] auth: valid session for GET /launcher +2026/08/22 11:24:09 server.go:393: [DEBUG] [web] ServeHTTP: GET /launcher from 172.18.0.4:60064 +2026/08/22 11:24:09 auth.go:134: [DEBUG] [web] auth: valid session for GET /backups/restore/app +2026/08/22 11:24:09 server.go:393: [DEBUG] [web] ServeHTTP: GET /backups/restore/app from 172.18.0.4:60064 +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true) +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3555MB avail≈22343MB +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26497444 → total=25898MB avail=22343MB used=3555MB (13.7%) +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.7GB used=8.0GB avail=57.2GB (11.6%) +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.5GB avail=884.7GB (0.6%) +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readLoadAvg: raw="0.99 1.36 1.08 3/810 5442" → 1m=0.99 5m=1.36 15m=1.08 +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 53.9°C (hwmon2) +2026/08/22 11:24:11 info.go:13: [DEBUG] [system] GetInfo done in 61ms — mem=3555MB/25898MB (13.7%), rootDisk=8.0GB/68.7GB (11.6%), load=0.99/1.36/1.08, temp=53.9°C (hwmon2), cpu=1.4% +2026/08/22 11:24:12 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan +2026/08/22 11:24:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/08/22 11:24:12 scheduler.go:67: [DEBUG] [scheduler] job stack-scan: execution starting +2026/08/22 11:24:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/08/22 11:24:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/08/22 11:24:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=true composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "cloudflared" deployed=false composePath=/opt/docker/stacks/cloudflared/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=true composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/08/22 11:24:12 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health +2026/08/22 11:24:12 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=true composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=true composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=true composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=false composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/08/22 11:24:12 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/08/22 11:24:12 manager.go:1497: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/08/22 11:24:12 manager.go:539: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/08/22 11:24:12 manager.go:570: [INFO] [stacks] ScanStacks complete: 56 stacks found (6 deployed, 50 available) +2026/08/22 11:24:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:24:12 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 47ms) +2026/08/22 11:24:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping romm — last check 2m50s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping calibre-web — last check 40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping kimai — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping opengist — last check 3m40s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping privatebin — last check 3m10s ago, effective interval 5m0s, healthy=true +2026/08/22 11:24:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (5 skipped not due, 1 skipped no container) +2026/08/22 11:24:12 manager.go:621: [INFO] [stacks] Status refresh: 14 containers across 56 stacks +2026/08/22 11:24:12 client.go:1142: [DEBUG] [agentapi] GET /storage -> 200 (559ms) +2026/08/22 11:24:12 scheduler.go:363: [INFO] [scheduler] Job agent-channel-health completed (took 559ms) +2026/08/22 11:24:13 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (660ms) +2026/08/22 11:24:14 client.go:1142: [DEBUG] [agentapi] GET /disks -> 200 (679ms) diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index b41618d7..1eb5cf66 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,7 @@ --- +| **R-356** | **The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no.** Shipped in controller v0.219.0. Evidence: `audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. **Reasoning kept:** *the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (`Manager.GetAppDrivePath`, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong.* **An app with no drive is not misconfigured** — `01-topology-and-trust.md` §8 carries the `[DESIGN]` marker; between 19 and 22 August that design was called a defect four times. **Deployment is asked of `ListDeployedStacks()` and FAILS CLOSED on a nil provider:** "cannot tell" must not become "go ahead" when the caller's next act is a write. **Two different failures get two different sentences** — installed-but-no-resolvable-data-root has its own refusal and its own route; widening `nincs telepítve` to cover it would send a customer to reinstall a running app and hide the real fault. **Measured, and load-bearing: 53 templates, 13 `needs_hdd: true`, 40 `false`** (catalogue @ `459766cb1639`). **The capture side's raw `GetStackHDDPath` is FENCED and was not changed** — capture resolves an app's declared `userdata`/`import` file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.219.0, 2026-08-22; `privatebin` on `demo-hp`: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: `git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` | | **R-216** | **A correct recovery code was reported to the customer as wrong.** Shipped in 0.120.0, v0.125.0. | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** Shipped in v0.203.0. Evidence: `documentation/tests/part4-rewalk-2026-08-06/journal.md`. | **CLOSED 2026-08-06 — controller v0.203.0, proven live.** *(State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.)* **The over-claim, as it stood: the fix covered the DECLARATION half only.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | | **R-219** | **The listing the screen promises could never render on the shape it exists for.** | **SHIPPED** (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index f6802c53..e3d7ef6a 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -525,7 +525,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-349** | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | | **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes | | **R-353** | **A restore reported success having returned configuration and no data — and no screen could have told the customer.** `demo-hp`, 2026-08-21, OpenGist. The off-site reconstitution refused at 16:37:14 (not installed); the person reinstalled and ran the local unit restore, which reported `Restore-from-unit completed: opengist in 8.666896042s`. **The unit it restored from contains `manifest.json` + `compose/{app.yaml,.felhom.yml,docker-compose.yml}` and NOTHING else — `volume_dumps: None`, `db_dumps: None`** — and the off-site snapshot was **182.3 KB**. So the restore returned the app's configuration; there was no data leg in the unit to return, and the outcome said only that it had completed. **A warning beside a success is read as a success, and an unknown must never be drawn as healthy.** **Compounding, and recorded as UNKNOWN rather than fine:** whether the 40-class reaches the off-site tier at all has **not been observed** — `runVolumeDumps` (`backup/backup.go:607+`) covers them on paper, but every unit on the box reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly dump run had happened on a one-hour-old box. | **OPEN — NEXT SESSION'S FIRST ITEM** | — | **Two things, in order. (1)** A restore whose unit carries no `db_dumps` and no `volume_dumps` must **say so in its outcome** — „a mentés csak a beállításokat tartalmazta, adatot nem" — instead of reporting a bare completion. The verdict must consult what was actually placed, not merely that the operation ended. **(2)** Then *prove* the off-site coverage of a named-volume app by running a dump cycle and reading the resulting manifest, rather than inferring it from the gate order. Do not close (1) on the strength of (2) being likely. **(2) IS NOW SATISFIED — drill 2026-08-21.** A dump cycle was run and the manifests read: `privatebin volume_dumps=[privatebin_privatebin_data.tar]`, `opengist volume_dumps=[opengist_opengist_data.tar]`, `kimai volume_dumps=[kimai_kimai_db_data.tar, kimai_kimai_var.tar]` — the 40-class DOES reach the off-site tier, and PrivateBin's planted 1 MB came back byte-identical from its off-site snapshot into the checking folder. **(1) stands and is now strictly larger than when written:** R-354 shows the bare completion is also reported over a unit that DID carry a data leg, because the off-site restore never replays volume dumps at all. | CC | -| **R-356** | **The off-site restore refuses for all 40 no-drive apps, says the app "is not installed" when it is running, and then gives an instruction those apps make impossible.** `ReconstituteFromOffsite` refuses when `GetStackHDDPath(stack)` is empty (`offbox_reconstitute.go:208-227`); for a 40-class app that is ALWAYS empty, because they are offered no storage field at deploy time (R-352's own measurement). Observed 2026-08-21 22:21 on `privatebin` while it was `deployed=true, state=running, healthy`: „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást ugyanerre a helyre…". **The predicate is "has an HDD path"; the sentence says "is not installed"; for this class they are different things**, and the remedy offered cannot be carried out. This is also what the 2026-08-21 afternoon OpenGist journey hit before falling back to the local restore (R-353). | **OPEN — HIGH** | — | Separate the two questions. A 40-class app has a destination — the system data path — and the restore already knows it. **RE-CHECKED 2026-08-22 against the corrected placement framing: this row SURVIVES UNCHANGED and is strengthened by it.** The 40-class has no `HDD_PATH` *because the architecture puts its hot data inside the guest by design* (`01-topology-and-trust.md:150-152`) — so an absent HDD path is the NORMAL state for most of the catalogue, and a restore that reads it as "the app is not installed" is misreading a correct configuration, not reporting a misconfiguration. The one sentence in this row that leaned on the old framing — "because they are offered no storage field at deploy time (R-352's own measurement)" — should read: *because these apps have no bulk volume to place, so the field correctly does not exist.* | CC | | **R-357** | **The DESTRUCTIVE restore has no free-space gate; the three that exist are all on non-destructive paths.** `offbox_reconstitute.go` contains **zero** references to `offboxFree`; the gates sit at `offbox_restore.go:231` (scratch restore), `:297` (prepare-full) and `:423` (place-to-live). Proven 2026-08-21 23:11 with 300 KB free and 1 MB to write: it stopped `paperless-ngx`, failed halfway (`rsync … No space left on device (28)`), left the data directory holding **2 of 5** planted entries, and restarted the app. The message is honest but is raw rsync output. | **OPEN — MEDIUM** | — | Same gate, same wording as `:297`, before the stop. | CC | | **R-358** | **A FAILED scratch restore leaves a partial copy that the product then offers as a full restore source — and the destructive restore runs from it and reports success.** `OffboxFullScratchReady` (`offbox_restore.go:305`) asks only whether the directory exists and is non-empty; its comment defers completeness to `PlaceOffsiteRestore`, which stats top-level placements, not files. Proven 2026-08-21 22:54-22:56 against a deliberately corrupted store: the restore failed honestly (`ciphertext verification failed`, 54 files, 15 of 16 originals), the wizard then offered „Teljes visszaállítás indítása", and pressing it reported `ok=true`. **The failure is detected and then forgotten.** | **OPEN — MEDIUM** | — | Record the failure against the scratch and refuse to place from it until it is re-prepared. | CC | | **R-359** | **The off-site restic store is never verified by anything, ever.** The complete set of restic verbs in the controller is `restore, snapshots, backup, unlock, stats, init, forget, prune, cat` — **no `check`**. The agent's `RestoreTest` is PBS-tier only. Established 2026-08-21 by deliberately corrupting one pack: `restic check` catches it immediately („ciphertext verification failed", „Fatal: repository contains errors"), and the product only meets the damage when a customer is already trying to recover. | **OPEN — MEDIUM** | — | A periodic `restic check` (structure) with an occasional `--read-data`, reported like any other backup verdict. Note PBS already has verify jobs; this is the tier that does not. | CC |