STATUS, report, register (R-691 closed live, R-704..R-706), Part E + D2 evidence
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+4
-1
@@ -18,7 +18,10 @@
|
||||
> **2026-09-28 — decided by CC unattended (operator may reverse): `09` §3 decision 43 — claper → PostgreSQL 17, calcom → 18** (by
|
||||
> each upstream; decision 42's rule; calcom's memory 768M → 1536M first, R-703 closed). Also that day: golden 0.276.0 baked + vouched with agent 0.137.0; controller
|
||||
> v0.277.0 (kept data loads from the off-site copy, R-691 (2)); floor 0.277.0; R-702 (claper default admin, P1) and
|
||||
> R-703 (calcom OOM at 768 MB) filed.
|
||||
> R-703 (calcom OOM at 768 MB, closed by `9555e73`) filed. **Afternoon (operator: finish today):** controller v0.278.0
|
||||
> (R-704 — a hold from an earlier install no longer holds a new one; floor 0.278.0); R-691 (2) proven LIVE on demo-hp
|
||||
> (full off-site restore + "use my kept data" from the off-site copy); night 27/28 read instead of a new night; R-705,
|
||||
> R-706 filed. Report: `REPORT-golden-kept-offsite-2026-09-28.md`.
|
||||
|
||||
## 2026-09-27 evening — a restore and a drive move keep the app's records (controller v0.276.0)
|
||||
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
# REPORT — 2026-09-28: weekly golden + Day-0 vouch, kept data from the off-site copy, claper/calcom, demo-hp space, night read, first live off-site restores
|
||||
|
||||
Brief: "the weekly golden, the Day-0 vouch, use my kept data from the off-site copy, two more PostgreSQL fixtures,
|
||||
demo-hp's restore-test space, the night watch, and the first live off-site restore" (revised 2026-09-28).
|
||||
Architecture read before any claim: `07-backup-architecture.md` §6.5, §6.6; `06-offsite-connectivity.md`;
|
||||
`03-host-agent.md` (restore-test storage); `09-update-architecture.md` §3 decisions 35–42.
|
||||
**Operator change mid-session (15:14):** "finish today" → Part E and D2 were done in the day, not overnight (below).
|
||||
|
||||
## The Parts
|
||||
|
||||
| Part | Step | State | Note |
|
||||
|---|---|---|---|
|
||||
| A | 1 read the manifest (rollback) | done | agent 0.132.0, golden 0.258.0, min_agent 0.131.0 |
|
||||
| A | 2 bake golden 0.276.0 | done, **with a deviation** | attempt 1 picked the `arm64` template; stopped by CC before `pct create` finished, nothing published. Attempt 2: all markers, leak grep 0 (control 1), 404 → 200, anonymous download sha equal, VM back to `virgin` |
|
||||
| A | 3 vouch (3 fields) | done | R-120 gate refused golden 0.258.0 first (negative control); `agent=0.137.0 golden=0.276.0 min_agent="0.131.0"` read back |
|
||||
| A | 4 Day-0 test install | done | installer 1.28.0, agent + golden fetched through the manifest and sha-verified, controller 0.276.0 healthy, claim gate armed; not claimed; scratch customer deleted by the hub's own cascade |
|
||||
| A | 5 golden-currency gate | done | green, not waived (record `documentation/tests/golden-0.276.0-2026-09-28/`); waiver file left as it is (to 2026-10-04) |
|
||||
| A | 6 agent CHANGELOG line | done | felhom-agent `5c68c86`; no floor move in Part A |
|
||||
| B | 1–3 code, tests, red-proofs | done | controller **v0.277.0**; RP1–RP3 |
|
||||
| B | 4 live Tier 1/Tier 2 regression on 9202 | Tier 1 done; **Tier 2 not runnable** | 9202 has one drive |
|
||||
| B | 5 release + floor | done | floor 0.277.0 at 10:11 (before 01:30), both boxes healthy |
|
||||
| B | — | **changed: a second release, v0.278.0** | R-704 blocked Part E; floor 0.278.0 at 15:36 |
|
||||
| C | calcom | done | memory fix first (R-703, 768M OOM at every start → 1536M); PG 16 → **18** proven on bench + box; catalog `037f956` |
|
||||
| C | claper | done | PG 16 → **17** proven on bench + box; catalog `4a249b9`; found R-702 (default admin) |
|
||||
| D | 1 R-701 (a) | done — **not enough** | 20.3 → 26.6 GiB free; 31 needed; 14:13 cycle refused again |
|
||||
| D | 2 night read | **changed** | the night 27/28 was read (logs, hub, agent journals), not tonight's; demo-felhom's off-site/update legs not readable |
|
||||
| E | 1 setup | done | nextcloud on demo-hp, seeded, joined the off-site copy |
|
||||
| E | 2 (a) off-site restore | done — **changed timing** | the snapshot came from the page's "run now" (15:40), not the night; seed back, marker gone |
|
||||
| E | 2 (b) kept data from off-site | done | choice named "távoli mentés, 2026-09-28 15:40"; seed + files back |
|
||||
| E | 3 teardown | done | app removed with its data; app list equal to before; verification copy deleted by the product |
|
||||
|
||||
## Claims in the brief that turned out wrong (or right)
|
||||
|
||||
- **"0.276.0's MinAgent is 0.131.0"** — right.
|
||||
- **"The R-120 gate accepts golden 0.276.0"** — right (it refused 0.258.0 and accepted 0.276.0).
|
||||
- **"The Day-0 test install leaves no hub record"** — **wrong.** It left a host, a vaulted break-glass credential, a claim code, a WireGuard peer (10.77.0.5, synced toward ep0) and reports. The customer DELETE cascade removed the hub side; the ep0 peer removal was **not observed** (ep0 untouched by rule).
|
||||
- **"The pairing code is shown"** — **wrong for this path.** The one-liner install shows no pairing code (that is the ISO path); what shows is the controller's claim gate (`dashboard not yet claimed`).
|
||||
- **"`pct fstrim` frees enough for R-701"** — **wrong.** It freed 6.3 GiB; the pool is 53.9 GiB and the guest holds ~26 GiB, so 31 GiB free cannot be reached by trimming.
|
||||
- **"paperless-ngx runs on demo-hp"** — right; it converted to PostgreSQL 18 in the night 27/28 (04:22 CEST, rows equal, 0 documents).
|
||||
- **"An app joins the off-site copy by a per-app switch"** — right (`POST /backup/offbox/toggle`, `app_backup.<app>.offbox`).
|
||||
|
||||
## Evidence
|
||||
|
||||
- Golden, vouch, Day-0, D1, D2: `documentation/audits/evidence-golden-0276-2026-09-28/`
|
||||
- Part B + E: `documentation/audits/kept-offsite-2026-09-28/` (redproofs/, live/, E/)
|
||||
- Part C: `documentation/audits/pg-calcom-claper-2026-09-28/`
|
||||
|
||||
## Rows
|
||||
|
||||
Opened: R-702 (claper default admin, P1), R-703 (calcom OOM — closed the same day), R-704 (leftover holds — fixed in
|
||||
0.278.0, WATCHING), R-705 (no "run the night now"), R-706 (verification copy survives removal). Closed: R-691, R-703.
|
||||
Updated: R-463, R-687, R-701. Register 337 → 342 rows. `unproven.py --summary`: not walked 35 of 55 (unchanged; the
|
||||
capability map was not edited).
|
||||
|
||||
## Teardown, three layers
|
||||
|
||||
- **Machines:** drill VM reverted to `virgin` (qemu gone); bench LXC 9401 created and destroyed twice; nextcloud,
|
||||
calcom, claper removed through the product on 9201/9202; 9202 pointed back to the live catalog; drill catalog = live.
|
||||
- **Hosts:** demo-hp `local-lvm` trimmed (kept); template cache files removed; no storage added.
|
||||
- **Hub:** scratch customer `drill-g0276` deleted by its cascade; manifest now golden 0.276.0 / agent 0.137.0; floor
|
||||
0.278.0. nextcloud's off-site snapshots stay in demo-hp's repository (removal never touches off-site history, R-474).
|
||||
@@ -1,42 +1,30 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-27 evening. Both demo boxes run controller 0.276.0 and host agent 0.137.0. Hub 0.125.0.**
|
||||
**Updated 2026-09-28 afternoon. Both demo boxes run controller 0.278.0 and host agent 0.137.0. Hub 0.125.0. New installs get golden 0.276.0 with agent 0.137.0.**
|
||||
|
||||
**Decisions I took on my own** (you may reverse each):
|
||||
1. **paperless-ngx goes to PostgreSQL 18, tandoor to 17.** Each follows what its own makers ship. paperless's makers use 18. tandoor's use 16, and 17 is the newest its framework supports.
|
||||
2. **The file browser reads nextcloud's kept folder by joining its group.** Nothing on your disk changes. The folder stays read-only in the view.
|
||||
1. **claper goes to PostgreSQL 17, calcom to 18.** Each follows what its own makers run: claper's makers use 15, so 17 (no disk-layout change); calcom's makers use 18.
|
||||
2. **calcom gets 1536 MB of memory instead of 768 MB.** At 768 MB it was killed at every start, so it could never run. At 1536 MB its own use peaked at about half.
|
||||
|
||||
**What I did, and it worked.**
|
||||
- **A backup now always says truthfully which version its data belongs to.** Before, the label changed to "new version" minutes after an update, while the data inside stayed old. A restore in that window broke docmost: its database would not start. Now the backup keeps the old version's settings next to the old data. I tested it on the scratch box: docmost came back whole at the old version, and the normal update brought it forward again.
|
||||
- **The box tested this by itself in the night.** It updated docmost's database at 02:15. Two minutes later the backup kept the matching old settings. The restore the next morning worked.
|
||||
- **Two more apps move to a new database version: paperless-ngx and tandoor.** Each was tested twice: on a throwaway test machine and on the scratch box through the normal update. Every account came back.
|
||||
- **adventurelog moves to its new version.** It keeps its world-map download. If that file was cut off, the box sets it aside and downloads it again. Tested both ways.
|
||||
- **The HP box's restore test now tests real backups of guests that exist**, not the golden template file and not an old backup of a deleted guest (two more small agent releases, each found when I checked the box).
|
||||
- **Small fixes:** the file browser no longer re-creates an empty kept folder; the "Kept data" name follows the box's language; after a load, the app page no longer shows a password that does not work.
|
||||
|
||||
**Evening follow-up (controller 0.276.0).**
|
||||
- **Fixed: moving an app to another drive made the box forget which version the app must run.** The app then took the newest version at its next start, and skipped the tested steps. Found by reading the code, not on a box. Tested in code only: no test box has two drives.
|
||||
- **Fixed: after a restore, the box forgot the old database copy, so it stayed on disk forever.** A restore now keeps that note, and also keeps whether you had the app switched on.
|
||||
- **Seen in the night: the scratch box deleted paperless's old database copy by itself**, after a backup made by the new database version. This is correct.
|
||||
- **Found in the night: the HP box's full-guest restore test can never run.** It now picks the right backup, but the disk has 21 GiB free and the test needs 31 GiB. It refuses safely every 6 hours. Filed, with three options; nothing decided.
|
||||
- **Not built: "Use my kept data" from the off-site copy.** It would be a new way to put data back, and no test box has an off-site copy to prove it on.
|
||||
- **The weekly golden is built and vouched** (0.276.0, with host agent 0.137.0). A test install in the throwaway machine came up on it with the right versions. The test customer is removed from the hub.
|
||||
- **"Use my kept data" can now load the database from the off-site copy.** Proven on the HP box: the page named "the off-site copy, 2026-09-28 15:40", the app came back with its account and its files.
|
||||
- **The first real restore from the off-site copy worked.** Account back, a later change gone (as it must be), same version.
|
||||
- **Two more apps can move to a new database version: claper and calcom.** Each passed the throwaway test machine and the scratch box.
|
||||
- **One more fix (controller 0.278.0): a "stopped" mark from an app's earlier install no longer sticks to a new install.** On the HP box such a mark from 13 September made the backup skip a freshly installed app. That app had no real backup at all. This is my second controller release today; the rules say one. It blocked the off-site proof, and nothing in the product could clear the mark.
|
||||
- **The HP box's night:** the database step ran, paperless moved to PostgreSQL 18 by itself (all rows equal), then the full-system backup ran. Nothing had to wait for anything.
|
||||
|
||||
**What broke, or is not done.**
|
||||
- **Found and fixed today: an app installed minutes ago could be updated with no backup of its database.** The update trusted a backup that held only the app's settings. Now it backs up first.
|
||||
- **"Use my kept data" still cannot load from the off-site copy.** Filed.
|
||||
- **The file-browser group fix is tested in code only.** No nextcloud kept folder existed on the scratch box.
|
||||
- **A backup stores the name of each app version, not the app itself.** If a maker deletes an old version, a restore of it cannot start. Today all 42 versions the catalog names still exist. Filed with options; nothing decided.
|
||||
- **claper creates an admin account with the public password "claper" on every install.** Anyone who knows that could log in. No box runs claper now. I did not change it; you choose the fix (below).
|
||||
- **The HP box's full-system restore test still cannot run.** Freeing space gave 26.6 GB; it needs 31 GB. It refuses safely, and the hub sees each refusal.
|
||||
- **The demo-felhom box's off-site step of last night is not readable** (today's upgrades erased the logs). Not a fault seen; just not read.
|
||||
- **Small gaps filed:** removing an app with its backups leaves its 1 GB off-site check copy; there is no button to run the whole night now.
|
||||
|
||||
**Rows.** 3 opened, 6 closed. The list went from 339 to 336. Evening and night: 2 opened, 1 closed; now 337.
|
||||
|
||||
**Decision for you — D4: when the only backup holds data of an older app version, what does a restore bring back?**
|
||||
- **A — the older version with its own data; then the normal update climbs, one tested step at a time (I recommend this, and it is built).** Cost: after the restore the app runs an older version for a night or until someone presses Update. The page says so.
|
||||
- **B — restore into a temporary copy, update it there, then move the data in.** Cost: a new mechanism on customer data, and double the disk space during the restore. Same end state as A.
|
||||
- **If you do nothing:** A stays in force. Nobody is blocked.
|
||||
**Rows.** 5 opened, 2 closed. The list went from 337 to 342 rows.
|
||||
|
||||
**What needs you.**
|
||||
1. **D4** above.
|
||||
2. **The weekly golden bake is due** (last one a week ago, 21 controller releases behind). The brief said no golden, so I renewed the waiver for 7 days only (to 4 October). If you do nothing, the checks go red again on 4 October and nothing can be pushed to felhom.eu until a bake or a new waiver.
|
||||
3. **Vouch agent 0.137.0** for new installs (hub → Configs → Day-0 artifacts). If you do nothing, new boxes install 0.134.0 and get the newer one only by a signed update.
|
||||
4. **The image-copy question** (a maker deletes an old version): keep as is, or copy installed versions into our registry. If you do nothing, nothing changes.
|
||||
5. **From before:** Peti's box in the Claude project text; Cloudflare leftovers of Peti's domain; the old Storage Box `PBS-storage-1`. If you do nothing, they stay as they are.
|
||||
1. **The claper admin password:** (A) remove claper from the catalog until it is fixed (I recommend this), or (B) have the box change that password after install. If you do nothing, a new claper install keeps the public password.
|
||||
2. **The HP box's restore test:** (A) let the test use the big NVMe disk (I recommend this; it has 880 GB free), or (B) accept that this small box cannot test it. If you do nothing, it refuses every 6 hours.
|
||||
3. **D4** (what a restore brings back when the backup is older): A stays in force. Nobody is blocked.
|
||||
4. **The image-copy question** (a maker deletes an old version): if you do nothing, nothing changes.
|
||||
5. **From before:** Peti's box in the project text; Peti's Cloudflare leftovers; the old Storage Box. If you do nothing, they stay.
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
======== felhom-pve (guest 9201) — night 2026-09-27/28, controller log 00:25Z–07:30Z (UTC; CEST = +2)
|
||||
-- agent (host) whole-guest backups:
|
||||
Sep 28 07:37:08 demo-felhom felhom-agent[2880126]: time=2026-09-28T07:37:08.524+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=backup:felhom-backup
|
||||
Sep 28 07:37:50 demo-felhom felhom-agent[2880126]: time=2026-09-28T07:37:50.932+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-backup archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_28-07_36_39.tar.zst size_bytes=2868142737 uncovered_vol
|
||||
Sep 28 07:37:50 demo-felhom felhom-agent[2880126]: time=2026-09-28T07:37:50.932+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-backup job=backup-9201-1790573799307855304 archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_28-07_
|
||||
Sep 28 08:07:07 demo-felhom felhom-agent[2880126]: time=2026-09-28T08:07:07.696+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-backup archive=felhom-backup:backup/vzdump-lxc-9201-2026_09
|
||||
Sep 28 08:07:08 demo-felhom felhom-agent[2880126]: time=2026-09-28T08:07:08.522+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
|
||||
Sep 28 08:08:36 demo-felhom felhom-agent[2880126]: time=2026-09-28T08:08:36.439+02:00 level=INFO msg="audit: gate decision" class=guest_destroy host=demo-felhom-8363b5 guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign key_id="" non
|
||||
Sep 28 08:08:36 demo-felhom felhom-agent[2880126]: time=2026-09-28T08:08:36.439+02:00 level=INFO msg="gate decision" class=guest_destroy guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign
|
||||
Sep 28 08:08:42 demo-felhom felhom-agent[2880126]: time=2026-09-28T08:08:42.485+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_27-07_35_21.tar.zst duration_s=94.787284561
|
||||
======== demo-hp (guest 9201) — night 2026-09-27/28, controller log 00:25Z–07:30Z (UTC; CEST = +2)
|
||||
-- agent (host) whole-guest backups:
|
||||
Sep 28 02:13:18 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:18.401+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
|
||||
Sep 28 04:43:18 demo-hp felhom-agent[1370926]: time=2026-09-28T04:43:18.400+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=backup:local
|
||||
Sep 28 04:44:58 demo-hp felhom-agent[1370926]: time=2026-09-28T04:44:58.683+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_28-04_37_06.tar.zst size_bytes=8825181955 uncovered_volumes=2
|
||||
Sep 28 04:44:58 demo-hp felhom-agent[1370926]: time=2026-09-28T04:44:58.683+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790563025947283530 archive=local:backup/vzdump-lxc-9201-2026_09_28-04_37_06.tar.zst
|
||||
Sep 28 08:13:18 demo-hp felhom-agent[1370926]: time=2026-09-28T08:13:18.401+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
|
||||
======== felhom-pve 9201 — debug-ring.log, night 2026-09-28 00:25Z–07:30Z
|
||||
-rw-r--r-- 1 root root 632864 Sep 28 13:52 /var/lib/docker/volumes/felhom-controller-data/_data/data/debug-ring.log
|
||||
{"timestamp":"2026-09-28T09:27:04Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-28T09:27:04Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
======== demo-hp 9201 — debug-ring.log, night 2026-09-28 00:25Z–07:30Z
|
||||
-rw-r--r-- 1 root root 902237 Sep 28 13:52 /var/lib/docker/volumes/felhom-controller-data/_data/data/debug-ring.log
|
||||
{"timestamp":"2026-09-28T13:36:07Z","level":"DEBUG","message":"[report] BuildReport: complete — containers=23, health=ok, deployed=11, available=42, app_telemetry=12","source":"builder.go:201"}
|
||||
{"timestamp":"2026-09-28T13:36:07Z","level":"DEBUG","message":"[report] Push: url=https://hub.felhom.eu/2026/09/28 02:30:00 [INFO] Event from demo-felhom: db_dump_completed (info) — Adatbázis mentés elkészült
|
||||
2026/09/28 02:32:04 [INFO] Event from demo-hp: db_dump_completed (info) — Adatbázis mentés elkészült
|
||||
2026/09/28 03:30:00 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: adventurelog
|
||||
2026/09/28 03:30:00 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: bentopdf
|
||||
2026/09/28 03:30:01 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: bookstack
|
||||
2026/09/28 03:30:01 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: calibre-web
|
||||
2026/09/28 03:30:01 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: docmost
|
||||
2026/09/28 03:30:02 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: kimai
|
||||
2026/09/28 03:30:02 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: opengist
|
||||
2026/09/28 03:30:05 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: paperless-ngx
|
||||
2026/09/28 03:30:05 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: privatebin
|
||||
2026/09/28 03:30:06 [INFO] Event from demo-hp: crossdrive_completed (info) — Másodlagos mentés elkészült: romm
|
||||
2026/09/28 06:00:35 [INFO] Event from demo-felhom: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (33s, a mentett adatok 100%-át újraolvasva)
|
||||
|
||||
## SUMMARY — Part D2, read 2026-09-28 ~16:00 CEST for the night 2026-09-27/28 (not a later night: the operator asked
|
||||
## to finish in the day). Sources, and why three: both controllers restarted at 10:11 and 15:36 (floor 0.277.0, 0.278.0),
|
||||
## which dropped their docker logs, and the persisted debug ring starts after the restarts. So: (1) the agents' journals
|
||||
## (above), (2) the hub's event log (above), (3) the demo-hp controller lines copied at 10:05 CEST, BEFORE the restarts
|
||||
## (phaseD2-paperless-demo-hp.txt) and demo-hp's settings.json off-site record read at 09:30 CEST.
|
||||
##
|
||||
## demo-hp (W = 02:30 CEST):
|
||||
## 02:32 db_dump_completed (hub event) · 03:30 crossdrive_completed ×10 (hub events) · off-site last_run 02:18:38Z =
|
||||
## 04:18 CEST, status ok, 90 snapshots (settings) · update leg: paperless-ngx 16 → 18 04:22:44–04:23:54 CEST, "step
|
||||
## ended done after 70.1 s" (controller, copied before the restart) · whole-guest local backup 04:37:06 → completed
|
||||
## 04:44:58 (agent), i.e. AFTER the gate opened (W+2h = 04:30) and after the leg ended → the gate had nothing to wait
|
||||
## for. R-687 item (4) (the gate waiting on a running leg) did NOT occur again.
|
||||
## demo-felhom (W = 02:30 CEST):
|
||||
## 02:30 db_dump_completed (hub event) · whole-guest backup to felhom-backup 07:36:39 → 07:37:50 (agent) · 06:00 off-site
|
||||
## integrity ok (hub event) · off-site leg and update leg: NOT READABLE (controller log dropped by today's restarts; no
|
||||
## hub event is sent for a successful off-site run or an update-leg step). Recorded as not read, not as fine.
|
||||
## R-701 side finding: every refused restore test reaches the hub — "[WARN] host demo-hp-bb76ea restore-test FAILED …
|
||||
## needs 31.0 GiB free, has 21.1 GiB" at each host-report (every 15 min).
|
||||
@@ -0,0 +1,60 @@
|
||||
15:17:28 # kp offsite-run — 2026-09-28T15:17:24+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
|
||||
15:17:28 E2 manual off-site run: the backup page's run-now button, POST /backup/offbox/run
|
||||
15:17:31 apps before: adventurelog Up 11 hours (healthy);adventurelog-frontend Up 11 hours (healthy);adventurelog-postgres Up 11 hours (healthy);bentopdf Up 11 hours (healthy);bookstack Up 11 hours (healthy);bookstack-db Up 11 hours (healthy);calibre-web Up 11 hours (healthy);cloudflared Up 4 days;docmost Up 11 hours (healthy);docmost-postgres Up 11 hours (healthy);docmost-redis Up 11 hours (healthy);felhom-controller Up 5 hours (healthy);filebrowser Up 4 days (healthy);kimai Up 11 hours (healthy);kimai-db Up 11 hours (healthy);nextcloud Up 5 hours (healthy);nextcloud-db Up 5 hours (healthy);nextcloud-redis Up 5 hours (healthy);opengist Up 11 hours (healthy);paperless-postgres Up 11 hours (healthy);paperless-redis Up 11 hours (healthy);paperless-webserver Up 11 hours (healthy);privatebin Up 11 hours (healthy);romm Up 11 hours (healthy);romm-db Up 11 hours (healthy);romm-redis Up 11 hours (healthy);traefik Up
|
||||
15:17:34 POST /backup/offbox/run -> ('HTTP/2 302', '/backups/remote?flash=flash.offbox.run_started')
|
||||
15:17:34 +0s active=True phase=dump app=
|
||||
15:19:29 +116s active=True phase= app=calibre-web
|
||||
15:19:34 +121s active=True phase= app=docmost
|
||||
15:19:39 +126s active=True phase= app=kimai
|
||||
15:19:54 +141s active=True phase= app=paperless-ngx
|
||||
15:20:10 +156s active=True phase= app=opengist
|
||||
15:20:15 +161s active=True phase= app=romm
|
||||
15:20:25 +171s active=True phase= app=nextcloud
|
||||
15:20:45 +191s active=True phase= app=bentopdf
|
||||
15:20:50 +196s active=True phase= app=bookstack
|
||||
15:20:55 +201s active=True phase=retention app=
|
||||
15:21:40 +246s active=False phase= app=
|
||||
15:21:40 status after: {'status': 'ok', 'last_run': '2026-09-28T13:21:36Z', 'last_error': '', 'snapshots': 91, 'last_duration': '3m57s'}
|
||||
15:21:44 controller lines:
|
||||
2026/09/28 13:18:12 backup.go:845: [INFO] [backup] Volume dump: docmost/docmost_docmost_storage → 2.5 KB
|
||||
2026/09/28 13:18:15 backup.go:845: [INFO] [backup] Volume dump: docmost/docmost_docmost_postgres_data → 67.2 MB
|
||||
2026/09/28 13:18:16 backup.go:845: [INFO] [backup] Volume dump: docmost/docmost_docmost_redis_data → 51.9 MB
|
||||
2026/09/28 13:18:16 backup.go:956: [INFO] [backup] Restarting docmost after volume dump
|
||||
2026/09/28 13:18:28 backup.go:946: [INFO] [backup] Stopping kimai for safe volume dump
|
||||
2026/09/28 13:18:32 backup.go:845: [INFO] [backup] Volume dump: kimai/kimai_kimai_db_data → 154.5 MB
|
||||
2026/09/28 13:18:34 backup.go:845: [INFO] [backup] Volume dump: kimai/kimai_kimai_var → 50.4 MB
|
||||
2026/09/28 13:18:34 backup.go:956: [INFO] [backup] Restarting kimai after volume dump
|
||||
2026/09/28 13:18:40 backup.go:729: [WARN] [backup] Skipping volume dump for nextcloud — the app is HELD stopped (its restore point is preserved)
|
||||
2026/09/28 13:18:40 backup.go:946: [INFO] [backup] Stopping opengist for safe volume dump
|
||||
2026/09/28 13:18:41 backup.go:845: [INFO] [backup] Volume dump: opengist/opengist_opengist_data → 181.0 KB
|
||||
2026/09/28 13:18:41 backup.go:956: [INFO] [backup] Restarting opengist after volume dump
|
||||
2026/09/28 13:18:41 backup.go:946: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/09/28 13:18:50 backup.go:845: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 16.1 MB
|
||||
2026/09/28 13:18:51 backup.go:845: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 69.3 MB
|
||||
2026/09/28 13:18:52 backup.go:845: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 21.8 MB
|
||||
2026/09/28 13:18:52 backup.go:956: [INFO] [backup] Restarting paperless-ngx after volume dump
|
||||
2026/09/28 13:19:04 backup.go:946: [INFO] [backup] Stopping privatebin for safe volume dump
|
||||
2026/09/28 13:19:05 backup.go:845: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 2.0 MB
|
||||
2026/09/28 13:19:05 backup.go:956: [INFO] [backup] Restarting privatebin after volume dump
|
||||
2026/09/28 13:19:05 backup.go:946: [INFO] [backup] Stopping romm for safe volume dump
|
||||
2026/09/28 13:19:10 main.go:898: [WARN] [deadapp] OOM scan failed: exec docker inspect -f {{.Name}}|{{.State.OOMKilled}}|{{.State.StartedAt}} adventurelog adventurelog-frontend adventurelog-postgres bentopdf bookstack bookstack-db calibre-web docmost docmost-p
|
||||
2026/09/28 13:19:10 backup.go:845: [INFO] [backup] Volume dump: romm/romm_romm_config → 2.5 KB
|
||||
2026/09/28 13:19:12 backup.go:845: [INFO] [backup] Volume dump: romm/romm_romm_db_data → 156.2 MB
|
||||
2026/09/28 13:19:12 backup.go:845: [INFO] [backup] Volume dump: romm/romm_romm_redis_data → 24.3 MB
|
||||
2026/09/28 13:19:12 backup.go:956: [INFO] [backup] Restarting romm after volume dump
|
||||
2026/09/28 13:19:23 backup.go:659: [INFO] [backup] App-data backup completed: 7 databases (14.2 MB total), 9 volume dump(s) (1m49.492s)
|
||||
2026/09/28 13:19:23 offbox.go:969: [INFO] [offbox] pre-push dump leg completed in 1m49.523s — snapshot pair is coherent
|
||||
2026/09/28 13:19:33 offbox.go:1391: [INFO] [offbox] backed up calibre-web (/mnt/felhom-drives/hdd_1/backups/primary/calibre-web, 1 mandatory path(s))
|
||||
2026/09/28 13:19:39 offbox.go:1391: [INFO] [offbox] backed up docmost (/mnt/sys_drive/felhom-data/backups/primary/docmost, 0 mandatory path(s))
|
||||
2026/09/28 13:19:50 offbox.go:1391: [INFO] [offbox] backed up kimai (/mnt/sys_drive/felhom-data/backups/primary/kimai, 0 mandatory path(s))
|
||||
2026/09/28 13:20:06 offbox.go:1391: [INFO] [offbox] backed up paperless-ngx (/mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx, 1 mandatory path(s))
|
||||
2026/09/28 13:20:09 offbox.go:1391: [INFO] [offbox] backed up privatebin (/mnt/sys_drive/felhom-data/backups/primary/privatebin, 0 mandatory path(s))
|
||||
2026/09/28 13:20:12 offbox.go:1391: [INFO] [offbox] backed up opengist (/mnt/sys_drive/felhom-data/backups/primary/opengist, 0 mandatory path(s))
|
||||
2026/09/28 13:20:20 offbox.go:1391: [INFO] [offbox] backed up romm (/mnt/felhom-drives/hdd_1/backups/primary/romm, 0 mandatory path(s))
|
||||
2026/09/28 13:20:45 offbox.go:1388: [WARN] [offbox] backed up nextcloud (/mnt/felhom-drives/hdd_1/backups/primary/nextcloud, 1 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's dat
|
||||
2026/09/28 13:20:48 offbox.go:1388: [WARN] [offbox] backed up bentopdf (/mnt/sys_drive/felhom-data/backups/primary/bentopdf, 0 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's dat
|
||||
2026/09/28 13:20:53 offbox.go:1391: [INFO] [offbox] backed up bookstack (/mnt/sys_drive/felhom-data/backups/primary/bookstack, 0 mandatory path(s))
|
||||
2026/09/28 13:21:36 offbox.go:1158: [INFO] [offbox] backup OK: 10 app(s) backed up, 91 snapshot(s), 3m57s
|
||||
2026/09/28 13:21:36 offbox_progress.go:342: [INFO] [offbox] manual run progress reporting ended after 4m2s
|
||||
|
||||
15:21:47 apps after: adventurelog Up 3 minutes (healthy);adventurelog-frontend Up 3 minutes (healthy);adventurelog-postgres Up 4 minutes (healthy);bentopdf Up 11 hours (healthy);bookstack Up 3 minutes (healthy);bookstack-db Up 3 minutes (healthy);calibre-web Up 3 minutes (healthy);cloudflared Up 4 days;docmost Up 3 minutes (healthy);docmost-postgres Up 3 minutes (healthy);docmost-redis Up 3 minutes (healthy);felhom-controller Up 5 hours (healthy);filebrowser Up 4 days (healthy);kimai Up 3 minutes (healthy);kimai-db Up 3 minutes (healthy);nextcloud Up 5 hours (healthy);nextcloud-db Up 5 hours (healthy);nextcloud-redis Up 5 hours (healthy);opengist Up 3 minutes (healthy);paperless-postgres Up 2 minutes (healthy);paperless-redis Up 2 minutes (healthy);paperless-webserver Up 2 minutes (healthy);privatebin Up 2 minutes (healthy);romm Up 2 minutes (healthy);romm-db Up 2 minutes (healthy);romm-redis Up 2 minutes (h
|
||||
@@ -0,0 +1,56 @@
|
||||
2026/09/28 13:02:10 backup.go:1145: [INFO] [backup] Found 21 DB dump files across drives
|
||||
bc851f0168e0 nextcloud-redis nextcloud redis:7-alpine
|
||||
b1bbeca1b46d nextcloud-db nextcloud mariadb:12.3
|
||||
2026/09/28 13:02:10 dbdump.go:208: [INFO] [backup] Discovered 7 databases
|
||||
2026/09/28 13:02:10 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/28 13:02:10 [INFO] [backup] Discovered app data: 11 apps
|
||||
2026/09/28 13:07:09 backup.go:1145: [INFO] [backup] Found 21 DB dump files across drives
|
||||
bc851f0168e0 nextcloud-redis nextcloud redis:7-alpine
|
||||
b1bbeca1b46d nextcloud-db nextcloud mariadb:12.3
|
||||
2026/09/28 13:07:10 dbdump.go:208: [INFO] [backup] Discovered 7 databases
|
||||
2026/09/28 13:07:10 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/28 13:07:10 [INFO] [backup] Discovered app data: 11 apps
|
||||
2026/09/28 13:12:10 backup.go:1145: [INFO] [backup] Found 21 DB dump files across drives
|
||||
bc851f0168e0 nextcloud-redis nextcloud redis:7-alpine
|
||||
b1bbeca1b46d nextcloud-db nextcloud mariadb:12.3
|
||||
2026/09/28 13:12:10 dbdump.go:208: [INFO] [backup] Discovered 7 databases
|
||||
2026/09/28 13:12:10 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/28 13:12:10 [INFO] [backup] Discovered app data: 11 apps
|
||||
2026/09/28 13:17:09 backup.go:1145: [INFO] [backup] Found 21 DB dump files across drives
|
||||
bc851f0168e0 nextcloud-redis nextcloud redis:7-alpine
|
||||
b1bbeca1b46d nextcloud-db nextcloud mariadb:12.3
|
||||
2026/09/28 13:17:10 dbdump.go:208: [INFO] [backup] Discovered 7 databases
|
||||
2026/09/28 13:17:10 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/28 13:17:10 [INFO] [backup] Discovered app data: 11 apps
|
||||
bc851f0168e0 nextcloud-redis nextcloud redis:7-alpine
|
||||
b1bbeca1b46d nextcloud-db nextcloud mariadb:12.3
|
||||
2026/09/28 13:17:34 dbdump.go:208: [INFO] [backup] Discovered 7 databases
|
||||
2026/09/28 13:17:34 backup.go:571: [INFO] [backup] Discovered 7 database(s): nextcloud-db(mariadb), romm-db(mariadb), paperless-postgres(postgres), kimai-db(mariadb), docmost-postgres(postgres), bookstack-db(mariadb), adventurelog-postgres(postgres)
|
||||
2026/09/28 13:17:34 dbdump.go:410: [INFO] [backup] DB dump: nextcloud-db → nextcloud-mariadb.sql (852.7 KB, 458ms, 131 tables)
|
||||
2026/09/28 13:17:35 dbdump.go:410: [INFO] [backup] DB dump: romm-db → romm-mariadb.sql (113.0 KB, 310ms, 36 tables)
|
||||
-rw-r--r-- 1 root root 789 2026-09-28T13:17:34 /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/data-stamps.json
|
||||
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud:
|
||||
total 16
|
||||
drwxr-xr-x 3 root root 4096 2026-09-28T13:17:34 .
|
||||
drwxr-xr-x 6 root root 4096 2026-09-28T13:17:34 ..
|
||||
-rw-r--r-- 1 root root 789 2026-09-28T13:17:34 data-stamps.json
|
||||
drwxr-xr-x 2 root root 4096 2026-09-28T13:17:34 db-dumps
|
||||
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps:
|
||||
total 864
|
||||
drwxr-xr-x 2 root root 4096 2026-09-28T13:17:34 .
|
||||
drwxr-xr-x 3 root root 4096 2026-09-28T13:17:34 ..
|
||||
-rw-r--r-- 1 root root 873125 2026-09-28T13:17:34 nextcloud-mariadb.sql
|
||||
|
||||
2026/09/28 13:20:45 offbox.go:1388: [WARN] [offbox] backed up nextcloud (/mnt/felhom-drives/hdd_1/backups/primary/nextcloud, 1 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's data; the next run with a dump leg will replace it
|
||||
---
|
||||
2026/09/28 13:17:52 backup.go:736: [DEBUG] [backup] bentopdf has no named volumes — volume dump skipped
|
||||
2026/09/28 13:18:40 backup.go:729: [WARN] [backup] Skipping volume dump for nextcloud — the app is HELD stopped (its restore point is preserved)
|
||||
|
||||
state running | hold_reason A(z) nextcloud frissítése 2026-09-13 21:56-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-13 21:51 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
|
||||
/db_validations/nextcloud-mariadb.sql {"validated_at": "2026-09-28T13:17:34Z", "table_count": 131, "has_header": true, "size": 873125, "mod_time": "2026-09-28T13:17:34Z"}
|
||||
/db_validations/pre-restore-20260913T195126Z-nextcloud-mariadb.sql {"validated_at": "2026-09-13T19:53:33Z", "table_count": 131, "has_header": true, "size": 502631, "mod_time": "2026-09-13T19:51:26Z"}
|
||||
/app_backup/nextcloud {"enabled": false, "offbox": true}
|
||||
/restore_holds/nextcloud {"stack": "nextcloud", "at": "2026-09-13T19:56:32Z", "reason": "update_failed", "copy_date": "2026-09-13T19:51:13Z", "copy_tier": 1, "copy_holds": "csak a be\u00e1ll\u00edt\u00e1sokat \u00e9s az adatb\u00e1zist tartalmazza, a f\u00e1jlokat nem"}
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
stop 200
|
||||
remove (with drive data + backups) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": ["/mnt/felhom-drives/hdd_1/appdata/nextcloud (126M)"], "hdd_paths_preserved": [], "backup_paths_removed": ["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (868K)"], "verified": true}, "message": "Stack nextcloud rem
|
||||
after: deployed False hold_reason None
|
||||
restore_holds.nextcloud = None
|
||||
app_backup.nextcloud = None
|
||||
/mnt/felhom-drives/hdd_1/appdata:
|
||||
paperless
|
||||
romm
|
||||
|
||||
/mnt/felhom-drives/hdd_1/backups/primary:
|
||||
calibre-web
|
||||
paperless-ngx
|
||||
romm
|
||||
0
|
||||
2026/09/28 13:36:14 waiter.go:95: [INFO] [report] hub wait channel active (hold ≤240s)
|
||||
2026/09/28 13:37:49 settings.go:1935: [INFO] [settings] update_failed hold CLEARED for nextcloud (set 2026-09-13T19:56:32Z; its install is gone)
|
||||
2026/09/28 13:37:49 router.go:1031: [INFO] [api] remove nextcloud: its update hold is cleared with it (R-491)
|
||||
|
||||
@@ -0,0 +1,71 @@
|
||||
15:38:14 # kp deploy — 2026-09-28T15:38:11+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:38:17 before: deployed= False | appdata: paperless romm
|
||||
15:38:21 deploy (no choice) -> 202 {"ok": true, "message": "Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
15:39:36 deployed after (75.4, 'running')
|
||||
15:39:39 # kp seed — 2026-09-28T15:39:36+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:39:48 nextcloud: occ user:add :: The account "drillf2b061" was created successfully Display name set to "drillf2b061"
|
||||
15:39:51 nextcloud: seeded user drillf2b061
|
||||
15:39:51 seeded (the account): True uid= drillf2b061
|
||||
15:39:56 PUT before-backup.txt -> 201
|
||||
15:39:57 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
|
||||
15:39:57 POST /backup/offbox/toggle app=nextcloud enabled=on -> ('HTTP/2 302', '/backups/remote?flash=flash.offbox.app_setting_updated')
|
||||
15:40:00 # kp offsite-run — 2026-09-28T15:39:57+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:40:00 E2 manual off-site run: the backup page's run-now button, POST /backup/offbox/run
|
||||
15:40:03 apps before: adventurelog Up 22 minutes (healthy);adventurelog-frontend Up 22 minutes (healthy);adventurelog-postgres Up 22 minutes (healthy);bentopdf Up 11 hours (healthy);bookstack Up 21 minutes (healthy);bookstack-db Up 22 minutes (healthy);calibre-web Up 21 minutes (healthy);cloudflared Up 4 days;docmost Up 21 minutes (healthy);docmost-postgres Up 21 minutes (healthy);docmost-redis Up 21 minutes (healthy);felhom-controller Up 3 minutes (healthy);filebrowser Up 4 days (healthy);kimai Up 21 minutes (healthy);kimai-db Up 21 minutes (healthy);nextcloud Up About a minute (healthy);nextcloud-db Up About a minute (healthy);nextcloud-redis Up About a minute (healthy);opengist Up 21 minutes (healthy);paperless-postgres Up 21 minutes (healthy);paperless-redis Up 21 minutes (healthy);paperless-webserver Up 20 minutes (healthy);privatebin Up 20 minutes (healthy);romm Up 20 minutes (healthy);romm-db Up 20 min
|
||||
15:40:06 POST /backup/offbox/run -> ('HTTP/2 302', '/backups/remote?flash=flash.offbox.run_started')
|
||||
15:40:06 +0s active=True phase=dump app=
|
||||
15:42:12 +126s active=True phase= app=docmost
|
||||
15:42:22 +136s active=True phase= app=opengist
|
||||
15:42:27 +141s active=True phase= app=bentopdf
|
||||
15:42:32 +146s active=True phase= app=calibre-web
|
||||
15:42:37 +151s active=True phase= app=kimai
|
||||
15:42:47 +161s active=True phase= app=nextcloud
|
||||
15:43:47 +221s active=True phase= app=paperless-ngx
|
||||
15:43:52 +226s active=True phase= app=privatebin
|
||||
15:43:57 +231s active=True phase= app=romm
|
||||
15:44:02 +236s active=True phase=retention app=
|
||||
15:44:47 +281s active=False phase= app=
|
||||
15:44:47 status after: {'status': 'ok', 'last_run': '2026-09-28T13:44:44Z', 'last_error': '', 'snapshots': 91, 'last_duration': '4m32s'}
|
||||
15:44:51 controller lines:
|
||||
2026/09/28 13:41:16 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 761.6 MB
|
||||
2026/09/28 13:41:17 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 187.5 KB
|
||||
2026/09/28 13:41:17 backup.go:956: [INFO] [backup] Restarting nextcloud after volume dump
|
||||
2026/09/28 13:41:17 manager.go:1216: [INFO] [stacks] Starting stack: nextcloud
|
||||
2026/09/28 13:41:23 manager.go:1231: [INFO] [stacks] Stack nextcloud started successfully (took 6.4s)
|
||||
2026/09/28 13:41:23 backup.go:946: [INFO] [backup] Stopping opengist for safe volume dump
|
||||
2026/09/28 13:41:24 backup.go:845: [INFO] [backup] Volume dump: opengist/opengist_opengist_data → 181.0 KB
|
||||
2026/09/28 13:41:24 backup.go:956: [INFO] [backup] Restarting opengist after volume dump
|
||||
2026/09/28 13:41:25 backup.go:946: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/09/28 13:41:26 manager.go:1562: [INFO] [stacks] Stack nextcloud post-start status:
|
||||
2026/09/28 13:41:26 manager.go:1565: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
|
||||
2026/09/28 13:41:26 manager.go:1565: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 9 seconds (healthy)
|
||||
2026/09/28 13:41:26 manager.go:1565: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 9 seconds (healthy)
|
||||
2026/09/28 13:41:32 backup.go:845: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 16.1 MB
|
||||
2026/09/28 13:41:33 backup.go:845: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 69.3 MB
|
||||
2026/09/28 13:41:33 backup.go:845: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 21.8 MB
|
||||
2026/09/28 13:41:33 backup.go:956: [INFO] [backup] Restarting paperless-ngx after volume dump
|
||||
2026/09/28 13:41:45 backup.go:946: [INFO] [backup] Stopping privatebin for safe volume dump
|
||||
2026/09/28 13:41:45 backup.go:845: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 2.0 MB
|
||||
2026/09/28 13:41:45 backup.go:956: [INFO] [backup] Restarting privatebin after volume dump
|
||||
2026/09/28 13:41:46 backup.go:946: [INFO] [backup] Stopping romm for safe volume dump
|
||||
2026/09/28 13:41:54 backup.go:845: [INFO] [backup] Volume dump: romm/romm_romm_config → 2.5 KB
|
||||
2026/09/28 13:41:55 backup.go:845: [INFO] [backup] Volume dump: romm/romm_romm_db_data → 156.2 MB
|
||||
2026/09/28 13:41:55 backup.go:845: [INFO] [backup] Volume dump: romm/romm_romm_redis_data → 25.3 MB
|
||||
2026/09/28 13:41:55 backup.go:956: [INFO] [backup] Restarting romm after volume dump
|
||||
2026/09/28 13:42:06 backup.go:659: [INFO] [backup] App-data backup completed: 7 databases (14.2 MB total), 10 volume dump(s) (2m0.416s)
|
||||
2026/09/28 13:42:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/hdd_1/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/09/28 13:42:07 offbox.go:969: [INFO] [offbox] pre-push dump leg completed in 2m0.458s — snapshot pair is coherent
|
||||
2026/09/28 13:42:20 offbox.go:1391: [INFO] [offbox] backed up docmost (/mnt/sys_drive/felhom-data/backups/primary/docmost, 0 mandatory path(s))
|
||||
2026/09/28 13:42:23 offbox.go:1391: [INFO] [offbox] backed up opengist (/mnt/sys_drive/felhom-data/backups/primary/opengist, 0 mandatory path(s))
|
||||
2026/09/28 13:42:27 offbox.go:1388: [WARN] [offbox] backed up bentopdf (/mnt/sys_drive/felhom-data/backups/primary/bentopdf, 0 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's dat
|
||||
2026/09/28 13:42:31 offbox.go:1391: [INFO] [offbox] backed up bookstack (/mnt/sys_drive/felhom-data/backups/primary/bookstack, 0 mandatory path(s))
|
||||
2026/09/28 13:42:36 offbox.go:1391: [INFO] [offbox] backed up calibre-web (/mnt/felhom-drives/hdd_1/backups/primary/calibre-web, 1 mandatory path(s))
|
||||
2026/09/28 13:42:46 offbox.go:1391: [INFO] [offbox] backed up kimai (/mnt/sys_drive/felhom-data/backups/primary/kimai, 0 mandatory path(s))
|
||||
2026/09/28 13:43:46 offbox.go:1391: [INFO] [offbox] backed up nextcloud (/mnt/felhom-drives/hdd_1/backups/primary/nextcloud, 1 mandatory path(s))
|
||||
2026/09/28 13:43:52 offbox.go:1391: [INFO] [offbox] backed up paperless-ngx (/mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx, 1 mandatory path(s))
|
||||
2026/09/28 13:43:55 offbox.go:1391: [INFO] [offbox] backed up privatebin (/mnt/sys_drive/felhom-data/backups/primary/privatebin, 0 mandatory path(s))
|
||||
2026/09/28 13:44:02 offbox.go:1391: [INFO] [offbox] backed up romm (/mnt/felhom-drives/hdd_1/backups/primary/romm, 0 mandatory path(s))
|
||||
2026/09/28 13:44:44 offbox.go:1158: [INFO] [offbox] backup OK: 10 app(s) backed up, 91 snapshot(s), 4m32s
|
||||
2026/09/28 13:44:44 offbox_progress.go:342: [INFO] [offbox] manual run progress reporting ended after 4m38s
|
||||
|
||||
15:44:54 apps after: adventurelog Up 4 minutes (healthy);adventurelog-frontend Up 4 minutes (healthy);adventurelog-postgres Up 4 minutes (healthy);bentopdf Up 11 hours (healthy);bookstack Up 4 minutes (healthy);bookstack-db Up 4 minutes (healthy);calibre-web Up 4 minutes (healthy);cloudflared Up 4 days;docmost Up 3 minutes (healthy);docmost-postgres Up 4 minutes (healthy);docmost-redis Up 4 minutes (healthy);felhom-controller Up 8 minutes (healthy);filebrowser Up 4 days (healthy);kimai Up 3 minutes (healthy);kimai-db Up 3 minutes (healthy);nextcloud Up 3 minutes (healthy);nextcloud-db Up 3 minutes (healthy);nextcloud-redis Up 3 minutes (healthy);opengist Up 3 minutes (healthy);paperless-postgres Up 3 minutes (healthy);paperless-redis Up 3 minutes (healthy);paperless-webserver Up 3 minutes (healthy);privatebin Up 3 minutes (healthy);romm Up 2 minutes (healthy);romm-db Up 2 minutes (healthy);romm-redis Up 2 mi
|
||||
@@ -0,0 +1,50 @@
|
||||
15:45:06 # kp marker — 2026-09-28T15:45:03+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:45:10 occ user:add markeraa8543 :: The account "markeraa8543" was created successfully
|
||||
15:45:11 PUT after-snapshot.txt -> 201
|
||||
15:45:21 occ user:info — present ['drillf2b061', 'markeraa8543'], absent [] (+ a uid that cannot exist, the control): {'drillf2b061': True, 'markeraa8543': True, 'nobody6a6478': False}
|
||||
15:45:22 PROPFIND as the seeded user -> 207; present ['before-backup.txt', 'after-snapshot.txt']: [True, True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'after-snapshot.txt', 'before-backup.txt']
|
||||
15:45:25 # kp offsite-restore — 2026-09-28T15:45:22+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:45:34 1 prepare (mode=full, no confirm) -> ('HTTP/2 302', '/backups/restore/app?name=nextcloud&full_prep=nextcloud&full_size=1.0+GB')
|
||||
15:45:34 2 download (mode=full, confirm=1) -> ('HTTP/2 302', '/backups/restore/app?name=nextcloud&flash=flash.offbox.restore_started')
|
||||
15:45:34 download +0s restore-status running=True op=offbox-restore last.ok=True msg='A(z) homebox: 1 adatkötet visszaállítva — az alkalmazás újraindult.'
|
||||
15:46:10 download +36s restore-status running=False op=offbox-restore last.ok=True msg='A(z) nextcloud teljes mentése visszaállítva ellenőrző mappába: /mnt/felhom-drives/hdd_1/backups/offsite-restore/nextcloud — a saját fájljaiddal együtt. A meglévő adatok változatlanok.'
|
||||
15:46:10 3 reconstitute (confirm=1) -> HTTP/2 302 /backups/restore/app?name=nextcloud&flash=flash.offbox.full_restore_started
|
||||
15:46:10 reconstitute +0s restore-status running=True op=offbox-reconstitute last.ok=True msg='A(z) nextcloud teljes mentése visszaállítva ellenőrző mappába: /mnt/felhom-drives/hdd_1/backups/offsite-restore/nextcloud — a saját fájljaiddal együtt. A meglévő adatok változatlanok.'
|
||||
15:46:52 reconstitute +42s restore-status running=False op=offbox-reconstitute last.ok=True msg='A(z) nextcloud: 0 fájl és 3 adatkötet és az adatbázis visszaállítva (mentés: 2026-09-28 15:40) — az alkalmazás újraindult.'
|
||||
15:46:56 controller lines:
|
||||
2026/09/28 13:45:34 offbox_handlers.go:359: [INFO] [web] off-box full-restore prepared for nextcloud (size 1.0 GB) — awaiting the customer's confirm; no restore has started
|
||||
2026/09/28 13:46:08 offbox_restore.go:400: [INFO] [offbox] restored nextcloud (6cb379a8, full=true) → /mnt/felhom-drives/hdd_1/backups/offsite-restore/nextcloud
|
||||
2026/09/28 13:46:08 offbox_handlers.go:383: [INFO] [web] off-box restore nextcloud completed (full=true, async)
|
||||
2544a1e0dcaf nextcloud nextcloud nextcloud...
|
||||
2026/09/28 13:46:13 dbdump.go:410: [INFO] [backup] DB dump: nextcloud-db → pre-restore-20260928T134613Z-nextcloud-mariadb.sql (867.8 KB, 517ms, 131 tables)
|
||||
2026/09/28 13:46:13 offbox_reconstitute.go:262: [INFO] [offbox] nextcloud: pre-restore safety dump written → pre-restore-20260928T134613Z-nextcloud-mariadb.sql (867.8 KB)
|
||||
2026/09/28 13:46:13 manager.go:1305: [INFO] [stacks] Stopping stack: nextcloud
|
||||
2544a1e0dcaf nextcloud nextcloud nextcloud...
|
||||
2026/09/28 13:46:15 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/28 13:46:15 manager.go:1314: [INFO] [stacks] Stack nextcloud stopped successfully (took 1.7s)
|
||||
2026/09/28 13:46:15 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_db_data for nextcloud
|
||||
2026/09/28 13:46:17 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_html for nextcloud
|
||||
2026/09/28 13:46:31 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_redis_data for nextcloud
|
||||
2026/09/28 13:46:31 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for nextcloud
|
||||
2026/09/28 13:46:31 manager.go:1269: [INFO] [stacks] Starting stack nextcloud services only: [nextcloud-db]
|
||||
2026/09/28 13:46:32 manager.go:1280: [INFO] [stacks] Stack nextcloud services [nextcloud-db] started (took 0.4s)
|
||||
2026/09/28 13:46:32 restore_db.go:87: [INFO] [backup] Restore nextcloud: replaying DB dump into nextcloud-db (mariadb)
|
||||
2026/09/28 13:46:37 dbdump.go:801: [INFO] [backup] Imported DB dump nextcloud-mariadb.sql into nextcloud-db (mariadb)
|
||||
2026/09/28 13:46:37 restore_db.go:97: [INFO] [backup] Restore nextcloud: replayed 1 DB dump(s)
|
||||
2026/09/28 13:46:37 manager.go:1216: [INFO] [stacks] Starting stack: nextcloud
|
||||
2026/09/28 13:46:44 manager.go:1231: [INFO] [stacks] Stack nextcloud started successfully (took 6.1s)
|
||||
2026/09/28 13:46:47 manager.go:1562: [INFO] [stacks] Stack nextcloud post-start status:
|
||||
2026/09/28 13:46:47 manager.go:1565: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
|
||||
2026/09/28 13:46:47 manager.go:1565: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 15 seconds (healthy)
|
||||
2026/09/28 13:46:47 manager.go:1565: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 9 seconds (healthy)
|
||||
2026/09/28 13:46:52 offbox_reconstitute.go:934: [INFO] [offbox] reconstituted nextcloud from snapshot 6cb379a8: 0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed, safety dump=pre-restore-20260928T134613Z-nextcloud-mariadb.sql, skewed=false
|
||||
2026/09/28 13:46:52 offbox_handlers.go:472: [INFO] [web] off-box reconstitute nextcloud completed (async): files=0 dbs=1 snapshot=6cb379a8
|
||||
|
||||
15:46:56 after: state= running pinned= {'nextcloud': 'nextcloud:34.0.4-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'}
|
||||
15:47:08 # kp read no-marker — 2026-09-28T15:47:05+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:47:18 occ user:info — present ['drillf2b061'], absent ['markeraa8543'] (+ a uid that cannot exist, the control): {'drillf2b061': True, 'markeraa8543': False, 'nobody47dfb6': False}
|
||||
15:47:18 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
|
||||
15:47:22 state: running images: nextcloud=nextcloud:34.0.4-apache nextcloud-redis=redis:7-alpine nextcloud-db=mariadb:12.3
|
||||
on disk after (a): before-backup.txt
|
||||
after-snapshot.txt
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
15:47:40 # kp remove-keep delete-backups — 2026-09-28T15:47:37+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:47:42 stop 200
|
||||
15:48:25 remove (keep drive data, remove_backups=True) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": [], "hdd_paths_preserved": ["/mnt/felhom-drives/hdd_1/appdata/nextcloud (126M)"], "backup_paths_removed": ["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (929M)"], "verified": true}, "message": "Stack nextcloud rem
|
||||
15:48:36 after: deployed= False | appdata: nextcloud paperless romm
|
||||
ls: cannot access '/mnt/felhom-drives/hdd_1/backups/primary/nextcloud': No such file or directory
|
||||
ls: cannot access '/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps': No such file or directory
|
||||
ls: cannot access '/mnt/*/*/backups/secondary/nextcloud': No such file or directory
|
||||
ls: cannot access '/mnt/*/backups/secondary/nextcloud': No such file or directory
|
||||
|
||||
15:48:39 # kp kept — 2026-09-28T15:48:36+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:48:42 GET /kept-data: ...datok Az eltávolított alkalmazások lemezen maradt adatai. Megnézheted őket, betöltheted egy új telepítésbe, vagy törölheted. Magától a szerver soha nem töröl közülük semmit. Erről nem készül mentés. Nextcloud 2026-09-28 15:39 · 125.5 MB Visszatölthető innen: távoli mentés, 2026-09-28 15:40 Betöltés Megnézem Törlés A(z) Nextcloud megőrzött adatai (125.5 MB) véglegesen törlődnek. Ezt nem lehet visszacsinálni. A törléshez írd be: Nextcloud Végleges törlés (function(){ var burger=document.querySelector('.nav-burger'); var sidebar=document.getElementById('sidebar'); var backdrop=document.querySelector('.nav-backdrop'); if(!burger||!sidebar||!backdrop)return; function setOpen(open){ sidebar.classL...
|
||||
15:48:44 GET /kept-data?lang=en: ...ign out ↗ Back Kept data Data that removed apps left on the drives. You can look at it, load it into a new install, or delete it. The server never deletes any of it by itself. This is not backed up. Nextcloud 2026-09-28 15:39 · 125.5 MB Can be loaded from: the off-site copy, 2026-09-28 15:40 Load Look Delete The kept data of Nextcloud (125.5 MB) is deleted for good. This cannot be undone. To delete it, type: Nextcloud Delete for good (function(){ var burger=document.querySelector('.nav-burger'); var sidebar=document.getElementById('sidebar'); var backdrop=document.querySelector('.nav-backdrop'); if(!burger||!sidebar||!backdrop)return; function setOpen(open){ sidebar.classList.toggle('is-open...
|
||||
15:48:48 # kp ask — 2026-09-28T15:48:45+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:48:53 deploy, no choice, lang=hu -> 409 use_offered=True use_desc='Az adatbázist innen töltjük vissza: távoli mentés, 2026-09-28 15:40. A régi fájlokkal indítjuk az alkalmazást. Ami ezután változott, hiányozhat.' use_off=''
|
||||
15:48:59 deploy, no choice, lang=en -> 409 use_offered=True use_desc='We load the database from the off-site copy, 2026-09-28 15:40 and start the app with the old files. Changes made after that time can be missing.' use_off=''
|
||||
15:48:59 still not installed: False
|
||||
15:49:11 # kp use — 2026-09-28T15:49:08+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:49:20 deploy kept_data=use -> 202 {"ok": true, "message": "A(z) Nextcloud megőrzött adatainak betöltése elindult."}
|
||||
15:50:05 deployed after (45.3, 'degraded')
|
||||
15:50:43 controller lines:
|
||||
2026/09/28 13:49:20 kept_install.go:97: [INFO] [api] Deploy nextcloud: USE MY KEPT DATA — loading from off-site snapshot 6cb379a8 (tier 3, 2026-09-28T13:40:07Z) under the kept files [/mnt/felhom-drives/hdd_1/appdata/nextcloud]
|
||||
2026/09/28 13:49:20 kept_load.go:263: [INFO] [backup] kept load nextcloud: downloading the recovery unit alone from off-site snapshot 6cb379a8 (/mnt/felhom-drives/hdd_1/backups/primary/nextcloud)
|
||||
2026/09/28 13:49:39 restore_unit.go:538: [WARN] [backup] Restore nextcloud: generated replacement for [NEXTCLOUD_ADMIN_PASSWORD] — the credential was reset (old value unrecoverable); no data-encrypting key was involved, but a regenerated DATABASE password will not match the restored data directory's stored hash (R-12
|
||||
2026/09/28 13:49:39 restore_unit.go:393: [INFO] [backup] Restoring nextcloud from recovery unit /mnt/felhom-drives/hdd_1/backups/offsite-proof/nextcloud/mnt/felhom-drives/hdd_1/backups/primary/nextcloud: images=3, secrets recovered=2/3, data_keys=0
|
||||
2026/09/28 13:49:39 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_db_data for nextcloud
|
||||
2026/09/28 13:49:40 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_html for nextcloud
|
||||
2026/09/28 13:49:51 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_redis_data for nextcloud
|
||||
2026/09/28 13:49:52 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for nextcloud
|
||||
2026/09/28 13:49:52 deploy.go:718: [INFO] [stacks] nextcloud: the restore generated [NEXTCLOUD_ADMIN_PASSWORD] — the app's own login came back with its data; the page will not show the new value as the password
|
||||
2026/09/28 13:49:53 restore_db.go:87: [INFO] [backup] Restore nextcloud: replaying DB dump into nextcloud-db (mariadb)
|
||||
2026/09/28 13:50:00 restore_db.go:97: [INFO] [backup] Restore nextcloud: replayed 1 DB dump(s)
|
||||
2026/09/28 13:50:15 restore_unit.go:478: [INFO] [backup] Restore-from-unit completed: nextcloud — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
2026/09/28 13:50:15 kept_load.go:285: [INFO] [backup] kept load nextcloud from the off-site copy 6cb379a8 done in 55s (volumes 3/3, dbs 1/1)
|
||||
2026/09/28 13:50:15 kept.go:521: [INFO] [stacks] after_load nextcloud: nextcloud [php occ files:scan --all] in 689ms (err=<nil>): Starting scan for user 1 out of 2 (admin)
|
||||
|
||||
15:50:43 restore status: {"running": false, "op": "restore", "stack": "nextcloud", "started_at": "2026-09-28T13:49:20.088190642Z", "last": {"op": "restore", "stack": "nextcloud", "ok": true, "message": "A(z) Nextcloud a megőrzött adataival fut.", "finished_at": "2026-09-28T13:50:15.203177307Z"}, "last_recent": true}
|
||||
15:50:46 # kp read no-marker — 2026-09-28T15:50:43+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.278.0
|
||||
15:50:56 occ user:info — present ['drillf2b061'], absent ['markeraa8543'] (+ a uid that cannot exist, the control): {'drillf2b061': True, 'markeraa8543': False, 'nobodyd8af7d': False}
|
||||
15:50:57 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'after-snapshot.txt', 'before-backup.txt']
|
||||
15:50:59 state: running images: nextcloud=nextcloud:34.0.4-apache nextcloud-redis=redis:7-alpine nextcloud-db=mariadb:12.3
|
||||
@@ -0,0 +1,22 @@
|
||||
stop 200
|
||||
remove (with drive data + backups) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": ["/mnt/felhom-drives/hdd_1/appdata/nextcloud (126M)"], "hdd_paths_preserved": [], "backup_paths_removed": ["/mnt/felhom-drives/hdd_1/backups/primary/nextcloud (36K)"], "verified": true}, "message": "Stack nextcloud removed"}
|
||||
removed-empty-harness-dir
|
||||
/mnt/felhom-drives/hdd_1/appdata:
|
||||
paperless
|
||||
romm
|
||||
|
||||
/mnt/felhom-drives/hdd_1/backups/offsite-proof:
|
||||
|
||||
/mnt/felhom-drives/hdd_1/backups/offsite-restore:
|
||||
nextcloud
|
||||
|
||||
/mnt/felhom-drives/hdd_1/backups/primary:
|
||||
calibre-web
|
||||
paperless-ngx
|
||||
romm
|
||||
0
|
||||
app_backup.nextcloud None | hold None
|
||||
|
||||
## app list AFTER Part E (deployed): ['adventurelog', 'bentopdf', 'bookstack', 'calibre-web', 'docmost', 'kimai', 'opengist', 'paperless-ngx', 'privatebin', 'romm']
|
||||
15:52:31 POST /backup/offbox/verify-copy/delete stack=nextcloud confirm=1 -> ('HTTP/2 302', '/backups/restore?flash=flash.offbox.scratch_deleted')
|
||||
15:52:34 offsite-restore now:
|
||||
@@ -0,0 +1,20 @@
|
||||
# kept-offsite — 2026-09-28: "use my kept data" from the off-site copy (R-691 (2)), and the first live off-site restores
|
||||
|
||||
Architecture: `07-backup-architecture.md` §6.5 (kept data, now with the off-site copy), §6.6 (the data's version).
|
||||
Controller v0.277.0 (the feature) and v0.278.0 (R-704, found on the way). Method: endpoint level — the exact calls the
|
||||
pages make (`tools/kp.py`, `tools/walk.py`), the app's own front doors for data (nextcloud `occ`, WebDAV).
|
||||
|
||||
| file | what |
|
||||
|---|---|
|
||||
| `redproofs/RP1–RP3` | v0.277.0: local-only choice; no version refusal; copy removed after the outcome (first draft) |
|
||||
| `redproofs/RP4–RP6` | v0.278.0: old hold predicate; no drop on "use"; no drop on install |
|
||||
| `live/B4-9202-*` | Tier 1 route regression on 9202 (choice named "saját mentés"; seed + file back); teardown |
|
||||
| `live/B5-floor.txt`, `live/B6-floor-0278.txt` | floor 0.277.0, 0.278.0, both demo boxes |
|
||||
| `E/E0`, `E/E1` | demo-hp app list before; nextcloud installed, seeded, joined the off-site copy |
|
||||
| `E/E2`, `E/E2b` | the page's off-site run-now; the snapshot held NO data — a 2026-09-13 hold (R-704) |
|
||||
| `E/E3`, `E/E4` | old nextcloud removed (hold cleared); fresh install, seed, off-site run: snapshot `6cb379a8` with data |
|
||||
| `E/E5-a-offsite-restore.txt` | (a) prepare → download → reconstitute; seed back, marker user gone, same version |
|
||||
| `E/E6-b-kept-from-offsite.txt` | (b) remove keep data + delete local copies → choice names the off-site copy → use → seed + files back |
|
||||
| `E/E9-teardown.txt` | removed with data; app list equal; verification copy deleted (R-706) |
|
||||
|
||||
Seed passwords live in the session scratchpad only.
|
||||
@@ -0,0 +1,4 @@
|
||||
## 2026-09-28T13:36:03Z floor 0.277.0 -> 0.278.0 (min_agent 0.131.0 from the v0.278.0 header)
|
||||
HTTP/1.1 303 See Other
|
||||
Location: /configuration?flash=floor_set
|
||||
name="min_controller_version" value="0.278.0"
|
||||
@@ -0,0 +1,10 @@
|
||||
# RP4 — the pre-fix ClearUpdateHold predicate (update holds only)
|
||||
1930c1930
|
||||
< if !ok || (h.Reason != HoldReasonUpdateFailed && h.Reason != HoldReasonUnhealthyStop) {
|
||||
---
|
||||
> if !ok || h.Reason != HoldReasonUpdateFailed { // RED-PROOF RP4
|
||||
r704_hold_test.go:31: reason "unhealthy_stop": cleared=false still-held=true, want cleared=true
|
||||
--- FAIL: TestR704_ClearUpdateHoldClearsTheInstallsHoldsOnly (0.00s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.004s
|
||||
FAIL
|
||||
@@ -0,0 +1,10 @@
|
||||
# RP5 — the use-my-kept-data path does not drop a leftover hold
|
||||
98c98
|
||||
< r.dropLeftoverHold(name, "use-kept-data") // R-704: this is a new install of an app that is not installed
|
||||
---
|
||||
> // RED-PROOF RP5: no dropLeftoverHold
|
||||
kept_install_test.go:246: the earlier install's hold survived the new install: {Stack:cloudapp At:2026-09-13T19:56:32Z ReplayError: RollbackErr: SafetyDump: Reason:update_failed CopyDate: CopyTier:0 CopyHolds: UndoState: NoWholeCopy:false CopiesSeen:[] UnhealthyKind: Trip:0}
|
||||
--- FAIL: TestR704_AFreshInstallDropsALeftoverHold (0.01s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.021s
|
||||
FAIL
|
||||
@@ -0,0 +1,10 @@
|
||||
# RP6 — the plain install path does not drop a leftover hold
|
||||
500c500
|
||||
< r.dropLeftoverHold(name, "install")
|
||||
---
|
||||
> // RED-PROOF RP6: no dropLeftoverHold
|
||||
kept_install_test.go:287: deployStack must call dropLeftoverHold before DeployStack (drop=0 deploy=21792)
|
||||
--- FAIL: TestR704_AFreshInstallDropsALeftoverHold (0.02s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.028s
|
||||
FAIL
|
||||
@@ -97,6 +97,34 @@ def users_report(want_present, want_absent):
|
||||
return ok
|
||||
|
||||
|
||||
def form(path, fields):
|
||||
"""A FORM post the way the page sends it (session cookie + _csrf). Returns (status line, location)."""
|
||||
sess = open(f"{w.SC}/sess{os.getpid()}.txt").read().strip()
|
||||
csrf = open(f"{w.SC}/csrf{os.getpid()}.txt").read().strip()
|
||||
args = ["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}"]
|
||||
for k, v in fields.items():
|
||||
args += ["--data-urlencode", f"{k}={v}"]
|
||||
r = w.sh(args + [f"{w.BASE}{path}"], timeout=300)
|
||||
lines = (r.stdout or "").split("\n")
|
||||
loc = [l.split(":", 1)[1].strip() for l in lines if l.lower().startswith("location:")]
|
||||
return lines[0].strip() if lines else "?", (loc[0] if loc else "")
|
||||
|
||||
|
||||
def wait_restore(label, cap=3600):
|
||||
t0 = time.time(); last = None
|
||||
while time.time() - t0 < cap:
|
||||
d = w.ctl("GET", "/api/backup/restore-status")[1].get("data") or {}
|
||||
cur = (d.get("running"), d.get("op"), (d.get("last") or {}).get("ok"), (d.get("last") or {}).get("message"))
|
||||
if cur != last:
|
||||
say(f" {label} +{round(time.time()-t0)}s restore-status running={cur[0]} op={cur[1]} last.ok={cur[2]} msg={str(cur[3])[:220]!r}")
|
||||
last = cur
|
||||
if not d.get("running") and time.time() - t0 > 5:
|
||||
return d
|
||||
time.sleep(3)
|
||||
return {}
|
||||
|
||||
|
||||
w.login()
|
||||
say(f"# kp {step} {arg} — {time.strftime('%FT%T%z')}; guest {w.GUEST}; controller {g('cat /etc/felhom-controller-image').strip()}")
|
||||
|
||||
@@ -173,6 +201,43 @@ elif step == "kept":
|
||||
t = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ", h))
|
||||
i = t.find("extcloud")
|
||||
say(f"GET /kept-data{'?lang=en' if lang else ''}: ...{t[max(0, i-200):i+500]}..." if i >= 0 else f"GET /kept-data{'?lang=en' if lang else ''}: no nextcloud row")
|
||||
elif step == "snapshots":
|
||||
code, d = w.ctl("GET", "/backup/offbox/status")
|
||||
say("offbox status:", json.dumps(d, ensure_ascii=False)[:600])
|
||||
say(g("docker exec felhom-controller sh -c 'ls /opt/docker/felhom-controller/data/offbox' 2>&1 | tr '\\n' ' '"))
|
||||
say(g(f"docker logs --since {arg or '4h'} felhom-controller 2>&1 | grep -iE 'offbox|off-site|snapshot|nextcloud' | grep -v DEBUG | cut -c1-300 | tail -40"))
|
||||
elif step == "offsite-restore":
|
||||
since = g("date -u +%Y-%m-%dT%H:%M:%SZ").strip()
|
||||
say("1 prepare (mode=full, no confirm) ->", form("/backup/offbox/restore", {"app": APP, "mode": "full"}))
|
||||
say("2 download (mode=full, confirm=1) ->", form("/backup/offbox/restore", {"app": APP, "mode": "full", "confirm": "1"}))
|
||||
wait_restore("download")
|
||||
st, loc = form("/backup/offbox/reconstitute", {"app": APP, "confirm": "1"})
|
||||
say("3 reconstitute (confirm=1) ->", st, loc)
|
||||
if "ack_placement" in loc or "placement" in loc:
|
||||
say(" the page asked for the placement acknowledgement — sending it with confirm")
|
||||
say("3b reconstitute (confirm=1, ack_placement=1) ->", form("/backup/offbox/reconstitute", {"app": APP, "confirm": "1", "ack_placement": "1"}))
|
||||
wait_restore("reconstitute")
|
||||
say("controller lines:\n" + g(f"docker logs --since {since} felhom-controller 2>&1 | grep -iE 'offbox|reconstitut|restor|nextcloud|version' | grep -v DEBUG | cut -c1-320 | head -60"))
|
||||
st = w.stack(APP)
|
||||
say("after: state=", st.get("state"), "pinned=", (st.get("app_config") or {}).get("pinned_images"))
|
||||
elif step == "offsite-run":
|
||||
say("E2 manual off-site run: the backup page's run-now button, POST /backup/offbox/run")
|
||||
say("apps before:", g("docker ps --format '{{.Names}} {{.Status}}' | sort | tr '\\n' ';'")[:900])
|
||||
since = g("date -u +%Y-%m-%dT%H:%M:%SZ").strip()
|
||||
say("POST /backup/offbox/run ->", form("/backup/offbox/run", {}))
|
||||
t0 = time.time(); last = None
|
||||
while time.time() - t0 < 3600:
|
||||
c, d = w.ctl("GET", "/backup/offbox/status")
|
||||
p = (d.get("progress") or {}); cur = (p.get("active"), p.get("phase"), p.get("current_app"))
|
||||
if cur != last:
|
||||
say(f" +{round(time.time()-t0)}s active={cur[0]} phase={cur[1]} app={cur[2]}"); last = cur
|
||||
if not p.get("active") and time.time() - t0 > 20:
|
||||
break
|
||||
time.sleep(5)
|
||||
c, d = w.ctl("GET", "/backup/offbox/status")
|
||||
say("status after:", {k: d.get(k) for k in ("status", "last_run", "last_error", "snapshots", "last_duration")})
|
||||
say("controller lines:\n" + g(f"docker logs --since {since} felhom-controller 2>&1 | grep -iE 'offbox|nextcloud|snapshot|dump' | grep -v DEBUG | cut -c1-260 | tail -40"))
|
||||
say("apps after:", g("docker ps --format '{{.Names}} {{.Status}}' | sort | tr '\\n' ';'")[:900])
|
||||
elif step == "kept-delete":
|
||||
h = w.page("/kept-data")
|
||||
paths = sorted(set(re.findall(r'name="path" value="([^"]*nextcloud[^"]*)"', h)))
|
||||
|
||||
@@ -806,16 +806,18 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). | **WATCHING — P3; owner: CC (live off-site proof)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). **-- 2026-09-28 afternoon: LIVE-PROVEN on demo-hp 9201 (controller 0.278.0), endpoint level.** A throwaway nextcloud, seeded through its own front door, joined the off-site copy; the off-site run-now pushed snapshot `6cb379a8`. (a) The full off-site restore (prepare → download → reconstitute): 3 volumes + the database replayed, the seed read back, a marker user written after the snapshot read ABSENT, same versions. (b) Remove keeping the data + deleting the local copies → reinstall: the choice and the kept list both named "távoli mentés, 2026-09-28 15:40" / "the off-site copy, 2026-09-28 15:40"; "use my kept data" downloaded the unit alone and loaded 3/3 volumes + 1/1 database in 55 s; the seed and the kept files read back. App removed with its data; demo-hp's app list equals the list before. `audits/kept-offsite-2026-09-28/E/`. | **CLOSED — controller v0.277.0, live 2026-09-28** |
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` **-- 2026-09-28 14:13 CEST:** the first scheduled cycle after the trim was refused again ("needs 31.0 GiB free, has 23.9 GiB"). **Answered: each refusal DOES reach the hub** — `[WARN] host demo-hp-bb76ea restore-test FAILED …` at every host-report (every 15 min). | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
|
||||
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. | **OPEN — P1; owner: operator (fix shape), CC implements** |
|
||||
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. **-- 2026-09-28 later: FIXED in the catalog (`9555e73`, alone in its commit): memory 768M → 1536M.** Measured on 9202 through the drill catalog: at 2048M a 12-minute watch read `anon` steady ~780 MiB, 0 kills; at 1536M, sampled every 2 s from the container's birth, `anon` peaked at **817 MiB (53 %)** during start, `memory.peak` 1075 MiB, 0 kills, healthy. 1024M would sit at the 80 % `memory_tight` line. The seed route works at the new limit (`Calcom` fixture, catalog `b35fc7f`). `…/box/R703-0*.txt` | **CLOSED — catalog `9555e73`, 2026-09-28** |
|
||||
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. | **OPEN — P3; owner: CC** |
|
||||
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. **-- 2026-09-28 later: SECOND and worse instance, then FIXED in controller v0.278.0.** demo-hp's fresh nextcloud (installed 10:13) carried an UPDATE hold from a nextcloud of 2026-09-13 (set before v0.242.0 made removals clear update holds; nothing ever swept it). At the manual off-site run (15:17) the backup leg logged `Skipping volume dump for nextcloud — the app is HELD stopped`, captured no unit, and pushed a snapshot that `carried NO database dump and NO volume tar` — a freshly installed app silently NOT backed up. **Fix:** a removal also clears the crash-loop stop (`settings.ClearUpdateHold`), and a new install (plain or "use my kept data") drops a leftover update/crash-loop hold of an app that is not installed (`Router.dropLeftoverHold`); restore holds (R-379) untouched. Red-proofed RP4–RP6 (`audits/kept-offsite-2026-09-28/redproofs/`). Floor 0.278.0. **STILL OPEN: live proof of the install-time drop** (a box with a leftover hold on an uninstalled app). | **WATCHING — P2; owner: CC (install-time drop, live)** |
|
||||
| **R-705** | **[P3-LOW] There is no way to run the night's chain now — only its pieces.** Asked by the operator 2026-09-28 (to finish a proof in the day). What exists (read from source, controller v0.278.0): the backup page's off-site run-now (`POST /backup/offbox/run`) runs the dump leg first (the R-44 pre-phase: DB dumps, volume dumps with brief app stops, unit capture) and then the push — used live on demo-hp 2026-09-28 15:17, 3m57s; the debug API has `backup/dbdump`, `backup/crossdrive` (Tier 2), `backup/integrity`, `backup/offsite-proof`. **Missing:** the automatic update leg (`RunUpdateLeg`, chained only to the scheduled off-site job) and the whole-guest backup (the agent's, on its own 24 h / 7 d cadence) have no manual trigger; the only way to run the chain in order is to move the backup window (`POST /backups/window`), which takes W..W+2h at least. **Needs:** a debug action "run tonight's chain now" (dump → Tier 2 → off-site → update leg, in order, one at a time), and an agent-side "whole-guest backup now" for demo boxes. Not built. | **OPEN — P3; owner: CC** |
|
||||
| **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** Measured 2026-09-28 on demo-hp: after a full off-site restore of nextcloud (which leaves the downloaded copy in `backups/offsite-restore/nextcloud`, ~1 GB, by design, for the household to inspect), `POST /api/stacks/nextcloud/remove` with `remove_backups: true` removed the unit and listed `backup_paths_removed` WITHOUT the verification copy; it stayed until the restore page's own delete (`POST /backup/offbox/verify-copy/delete`, 302 `scratch_deleted`). A household that removes an app to free space keeps 1 GB it cannot see on the app list. **Fix direction:** the removal with backups also deletes the app's verification copy (the same `DeleteOffsiteRestoreCopy`). Evidence: `audits/kept-offsite-2026-09-28/E/E9-teardown.txt`. | **OPEN — P3; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user