docs: 09 decisions 35-38 + part 10 as built; R-689 filed, R-469 narrowed; night-2026-09-26 evidence (A, B, C, D, F)
gates / gates (push) Successful in 24s
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -402,6 +402,46 @@ R-636's louder repeated alarm.
|
||||
behaviour shipped in v0.271.0 is now the rule. **Why:** one night's delay for an app is cheap, and a resume
|
||||
would be one more mechanism acting with nobody watching.
|
||||
|
||||
### 2026-09-25 (evening) — two operator rulings
|
||||
|
||||
35. **Part 10 is built now, one app first** (operator ruling 2026-09-25 evening, D1 option A). The box
|
||||
converts a PostgreSQL major as a step of the guarded update. **docmost is the first app.** Each other app
|
||||
needs its own proof on both venues (the test bench and a scratch box through the real guarded Update)
|
||||
before the catalog may move it. The engine-major gate stays for every app without that proof. **Why:**
|
||||
decision 16 said "each of the eleven apps is proven before the catalog may move it"; one app first makes
|
||||
the first proof small enough to read. As built: §6.4 part 10.
|
||||
36. **A reinstall over kept data offers a choice** (operator ruling 2026-09-25 evening, D2 option A): "use
|
||||
my kept data" (the database from a backup, with the kept files) or "start fresh" (the kept files move to
|
||||
a dated folder; nothing is deleted). **And, from the operator's follow-up question: kept data must never
|
||||
be a dead end.** The household can see it (read only), load it into a later install, and delete it.
|
||||
Support is for the rare case, not the normal one. **Not ruled (D3, in `STATUS.md`):** whether the box
|
||||
ever deletes kept data by itself. Until it is ruled, **nothing deletes kept data automatically** — only a
|
||||
household's explicit Delete. Where it lives and who deletes it: `07-backup-architecture.md` §"Kept data".
|
||||
|
||||
### 2026-09-25 (evening) — two decisions taken by CC unattended, building part 10
|
||||
|
||||
37. **docmost converts PostgreSQL 16 → 18, not 16 → 17** — *decided by CC unattended 2026-09-25 — operator may
|
||||
reverse.* *One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount
|
||||
change; (b) 18 — the data volume's mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) every
|
||||
box converts again later, and docmost's own upstream compose already ships `postgres:18` at
|
||||
`/var/lib/postgresql`; (b) the step changes the mount point — measured: `postgres:18` REFUSES (exit 1) even an
|
||||
EMPTY volume at `/var/lib/postgresql/data` — so the step's definition carries the mount, and the undo must put
|
||||
the old mount back with the old bytes, which it does (it restores the old definition and every volume).
|
||||
**Why (b):** one conversion is better than two when the app's upstream runs the newer major. Reversible: the
|
||||
catalog could pin 17 instead before any customer box moves. `audits/night-2026-09-26/A/README.md` A1.
|
||||
38. **The box loads from a `pg_dumpall` of the OLD engine, with no error tolerated** — *decided by CC unattended
|
||||
2026-09-25 — operator may reverse.* *One sentence:* what does the conversion load into the new engine?
|
||||
**Options:** (a) the existing safety dump (`pg_dump` per database, `--no-owner --no-privileges`); (b) a
|
||||
`pg_dumpall` taken from the old engine at conversion time, loaded with `ON_ERROR_STOP`. **Costs:** (a) is free,
|
||||
and silently drops any role, grant or database setting beyond the bootstrap ones (for docmost both routes gave
|
||||
identical databases — it has one role and one database; the other ten are not measured); (b) one more dump
|
||||
(0.6 s for docmost) and a load that must step around exactly two objects the new engine's entrypoint makes (the
|
||||
bootstrap role and database — measured: loaded raw, exactly two `already exists` errors). **Why (b):** decision
|
||||
16 says "save everything"; the two collisions are removed precisely (the entrypoint's databases dropped only
|
||||
when they hold no table; `CREATE ROLE <x>;` skipped only for a role that exists — its `ALTER ROLE … PASSWORD`
|
||||
still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks
|
||||
(`adminpack`) stopped the load and the box undid it. `audits/night-2026-09-26/A/README.md` A2.
|
||||
|
||||
---
|
||||
|
||||
## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed
|
||||
@@ -1175,7 +1215,7 @@ what the part can do to a household's data if it is wrong, not how likely that i
|
||||
| **7** | **SHIPPED — controller v0.271.0 (`9cf13a3`), proven on 9202 over four simulated nights 2026-09-24/25 and watched through the demo boxes' first real automatic night (`audits/DRILL-night-2026-09-25.md`).** Decisions 31–33 taken unattended. Gaps a scratch box cannot close: R-687; no resume after a restart: R-686. *(Before:)* **SPIKED 2026-09-24 (night, Part C) — the measured build brief is §6.4.2 below.** **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has ENDED (on every exit path, skips included), one app at a time, one step per press, `app_update.unattended` default ON (decision 12), `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it, and the full-system backup's gate waits for it until W+5h (decision 20). | 11, 12, 14, 15, 20 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
|
||||
| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
|
||||
| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
|
||||
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
|
||||
| **10** | **SHIPPED FOR DOCMOST — controller v0.273.0 + catalog `6a4a5f0` (harness v4, the gate's proof clause), proven live on 9202 2026-09-25** (`audits/night-2026-09-26/D/`): one press converted docmost 16 → 18 in **9.9 s** of engine work (**42 s** end to end), the check equal (2 databases, 48 tables, 71 rows), `PG_VERSION` 18, the seed read back; the load failing (an extension 18 lacks) → **undone in 43.6 s**; the app unhealthy on 18 → **undone in 148 s** (90 s drill health timeout); the controller SIGKILLed one second after the volume was emptied → the restart **undid it in 40 s**; each ending on 16 with the seed read back. The old datadir's copy is kept until a backup is proven after the conversion. Decisions 35, 37, 38. Every other PostgreSQL app stays refused by the gate until its own two-venue proof. *(Before:)* **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
|
||||
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |
|
||||
|
||||
**Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
# A1 — what the postgres images declare (9202, 2026-09-25T10:20:36Z)
|
||||
## postgres:16-alpine
|
||||
digest=postgres@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea PGDATA/ENV=PG_MAJOR=16 PG_VERSION=16.15 PG_SHA256=c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed PGDATA=/var/lib/postgresql/data VOLUME={"/var/lib/postgresql/data":{}}
|
||||
## postgres:17-alpine
|
||||
digest=postgres@sha256:b0f9560a2de083e2cc7382e75f808c7381a32852a7ec49117deedb300e552b24 PGDATA/ENV=PG_MAJOR=17 PG_VERSION=17.11 PG_SHA256=dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979 PGDATA=/var/lib/postgresql/data VOLUME={"/var/lib/postgresql/data":{}}
|
||||
## postgres:18-alpine
|
||||
digest=postgres@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873 PGDATA/ENV=PG_MAJOR=18 PG_VERSION=18.6 PG_SHA256=555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f PGDATA=/var/lib/postgresql/18/docker VOLUME={"/var/lib/postgresql":{}}
|
||||
|
||||
# A1b — postgres:18-alpine with an EMPTY volume at /var/lib/postgresql/data (the templates' mount)
|
||||
state: exited exit=1
|
||||
Error: in 18+, these Docker images are configured to store database data in a
|
||||
format which is compatible with "pg_ctlcluster" (specifically, using
|
||||
major-version-specific directory names). This better reflects how
|
||||
PostgreSQL itself works, and how upgrades are to be performed.
|
||||
|
||||
See also https://github.com/docker-library/postgres/pull/1259
|
||||
|
||||
Counter to that, there appears to be PostgreSQL data in:
|
||||
/var/lib/postgresql/data (unused mount/volume)
|
||||
|
||||
This is usually the result of upgrading the Docker image without
|
||||
upgrading the underlying database using "pg_upgrade" (which requires both
|
||||
versions).
|
||||
|
||||
The suggested container configuration for 18+ is to place a single mount
|
||||
at /var/lib/postgresql which will then place PostgreSQL data in a
|
||||
subdirectory, allowing usage of "pg_upgrade --link" without mount point
|
||||
boundary issues.
|
||||
|
||||
See https://github.com/docker-library/postgres/issues/37 for a (long)
|
||||
discussion around this process, and suggestions for how to do so.
|
||||
|
||||
# A1c — control: postgres:18-alpine with an empty volume at /var/lib/postgresql (its own layout)
|
||||
state: running
|
||||
2026-09-25 10:21:17.475 UTC [1] LOG: listening on Unix socket "/var/run/postgresql/.s.PGSQL.5432"
|
||||
2026-09-25 10:21:17.488 UTC [66] LOG: database system was shut down at 2026-09-25 10:21:17 UTC
|
||||
2026-09-25 10:21:17.496 UTC [1] LOG: database system is ready to accept connections
|
||||
18
|
||||
@@ -0,0 +1,37 @@
|
||||
# A2/A3/A5 — docmost seeded on 9202, measured on a COPY of its datadir (2026-09-25T10:25:58Z)
|
||||
datadir size: 65.5M
|
||||
A3 counts (old engine): 0.49s, 48 tables, 62 rows total
|
||||
db:docmost owner=docmost enc=UTF8 coll=en_US.utf8
|
||||
db:postgres owner=docmost enc=UTF8 coll=en_US.utf8
|
||||
role:docmost super=true login=true pw=true
|
||||
ext:docmost:pg_trgm
|
||||
ext:docmost:plpgsql
|
||||
ext:docmost:unaccent
|
||||
A2 pg_dumpall: 0.59s, 132184 B, tail: -- PostgreSQL database dump complete
|
||||
A2 pg_dump --no-owner --no-privileges docmost: 0.27s, 123545 B
|
||||
what pg_dumpall carries that the per-db dump does not:
|
||||
1 CREATE ROLE docmost;
|
||||
1 CREATE DATABASE docmost WITH TEMPLATE = template0 ENCODING = 'UTF8' LOCALE_PROVIDER = libc LOCALE = 'en_US.utf8';
|
||||
1 ALTER TABLE public.workspaces OWNER TO docmost;
|
||||
1 ALTER TABLE public.workspace_invitations OWNER TO docmost;
|
||||
1 ALTER TABLE public.watchers OWNER TO docmost;
|
||||
1 ALTER TABLE public.users OWNER TO docmost;
|
||||
1 ALTER TABLE public.user_tokens OWNER TO docmost;
|
||||
1 ALTER TABLE public.user_sessions OWNER TO docmost;
|
||||
1 ALTER TABLE public.user_mfa OWNER TO docmost;
|
||||
1 ALTER TABLE public.templates OWNER TO docmost;
|
||||
1 ALTER TABLE public.spaces OWNER TO docmost;
|
||||
1 ALTER TABLE public.space_members OWNER TO docmost;
|
||||
## route dumpall: new engine init 2.64s, load 1.76s, counts 0.48s, PG_VERSION=18, load errors=2
|
||||
1 ERROR: database "docmost" already exists
|
||||
1 ERROR: role "docmost" already exists
|
||||
dbs/owners/extensions/row counts: IDENTICAL
|
||||
roles before: role:docmost super=true login=true pw=true
|
||||
roles after: role:docmost super=true login=true pw=true
|
||||
datadir after: 49.9M
|
||||
## route pgdump: new engine init 2.64s, load 1.44s, counts 0.49s, PG_VERSION=18, load errors=0
|
||||
dbs/owners/extensions/row counts: IDENTICAL
|
||||
roles before: role:docmost super=true login=true pw=true
|
||||
roles after: role:docmost super=true login=true pw=true
|
||||
datadir after: 49.8M
|
||||
A5 space: copy of the old datadir = 68707765 B; dump = 132184 B; new datadir = 52271397 B
|
||||
@@ -0,0 +1,13 @@
|
||||
# A4 — the undo after the DB volume was EMPTIED, by hand with the product's own helper commands (undo.go Copy/Restore), 9202 2026-09-25T10:27:34Z
|
||||
app stopped: 0 running
|
||||
copy with marker: ok
|
||||
volume emptied: 0 entries left
|
||||
Restore (the product's command): ok
|
||||
PG_VERSION after restore: 16
|
||||
2026-09-25 12:27:39.221 CEST [29] LOG: database system was shut down at 2026-09-25 12:27:35 CEST
|
||||
2026-09-25 12:27:39.230 CEST [1] LOG: database system is ready to accept connections
|
||||
copy removed
|
||||
12:34:29 docmost: login as the seeded user http=200 ok=True
|
||||
12:34:29 A4-after-undo-retry: read docmost=True
|
||||
12:34:30 docmost: login as the seeded user http=200 ok=True
|
||||
12:34:30 A4-after-undo-productstart: read docmost=True
|
||||
@@ -0,0 +1,88 @@
|
||||
# Part A — the PostgreSQL conversion spike on docmost (9202, 2026-09-25 midday)
|
||||
|
||||
Written BEFORE any build. docmost 0.96.0 / `postgres:16-alpine` installed on 9202 through the product
|
||||
(`POST /api/stacks/docmost/deploy`), seeded through its own API (a workspace + an account), read back
|
||||
(`data-readback.txt`: `A-read0`). Every measurement below ran on a COPY of its datadir.
|
||||
|
||||
## A1 — the target major: 18 (CC-unattended decision, `09` §3 decision 37)
|
||||
|
||||
**Measured (`A1-images.txt`):**
|
||||
|
||||
- `postgres:16-alpine` and `17-alpine`: `PGDATA=/var/lib/postgresql/data`, `VOLUME /var/lib/postgresql/data`.
|
||||
- `postgres:18-alpine` (18.6): `PGDATA=/var/lib/postgresql/18/docker`, `VOLUME /var/lib/postgresql`.
|
||||
- **18 with an EMPTY volume at `/var/lib/postgresql/data` (where all eleven templates mount it) REFUSES:
|
||||
exit 1** — *"in 18+, these Docker images are configured to store database data in a format which is
|
||||
compatible with pg_ctlcluster … there appears to be PostgreSQL data in: /var/lib/postgresql/data (unused
|
||||
mount/volume)"*. It refuses even when that mount is empty. Control: the same image with the volume at
|
||||
`/var/lib/postgresql` initialises and serves, `PG_VERSION` 18.
|
||||
- **docmost's upstream compose ships `image: postgres:18` with `db_data:/var/lib/postgresql`**
|
||||
(github docmost/docmost `docker-compose.yml`, read 2026-09-25). Its docs name no other version.
|
||||
|
||||
**So the brief's claim is TRUE**, and stronger than stated: 18 refuses at the old mount, even empty. A move to
|
||||
18 therefore needs the mount moved to `/var/lib/postgresql` in the same step, which the step's own definition
|
||||
carries (the template's compose / `steps/<key>.yml`).
|
||||
|
||||
**Decision 37 — docmost converts 16 → 18, not 16 → 17.**
|
||||
*One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount change;
|
||||
(b) 18 — the mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) is a second conversion
|
||||
later, and upstream already ships 18, so every box converts twice; (b) the step changes the volume's mount
|
||||
point, so an undo must put the OLD mount back with the old data — which the undo does anyway (it restores the
|
||||
old definition and the volume's bytes). **Why (b):** one conversion is better than two when the app's own
|
||||
upstream runs the newer major; the mount change rides the step's definition and the undo's existing
|
||||
definition restore. Reversible (the catalog can pin 17 instead before any box moves).
|
||||
|
||||
## A2 — what to load from: `pg_dumpall` from the old engine (decision 38)
|
||||
|
||||
`A2-A3-A5-measure.txt`. On the seeded datadir (48 tables, 62 rows, 49 MB):
|
||||
|
||||
| | size | time | what it carries |
|
||||
|---|---|---|---|
|
||||
| `pg_dumpall` (old engine) | 132 184 B | 0.59 s | roles WITH their password hashes, `CREATE DATABASE … LOCALE`, owners, grants |
|
||||
| `pg_dump --no-owner --no-privileges docmost` (today's safety dump) | 123 545 B | 0.27 s | one database's schema + data; no roles, no owners, no database settings |
|
||||
|
||||
Both loaded into a fresh 18 gave **identical** databases, owners, encodings, collations, extensions
|
||||
(`pg_trgm`, `plpgsql`, `unaccent`) and row counts for docmost — because docmost has ONE role (the bootstrap
|
||||
superuser the new engine's entrypoint recreates from the same env) and ONE database.
|
||||
|
||||
**So the brief's claim is TRUE** (the safety dump is per-database `pg_dump` with `--no-owner --no-privileges`
|
||||
— `appbackup/dbdump.go`), **and for docmost it would have been enough.** It would not be for an app with a
|
||||
second role or a second database, which the other ten have not been measured for.
|
||||
|
||||
**Decision 38 — the box loads from a `pg_dumpall` taken from the OLD engine at conversion time, and the load
|
||||
tolerates NO error.** *One sentence:* what does the box load into the new engine? **Options:** (a) the
|
||||
existing safety dump; (b) a `pg_dumpall` of the old engine, loaded with `ON_ERROR_STOP`. **Costs:** (a) is
|
||||
free, and silently drops any role, grant or database setting beyond the bootstrap ones; (b) one more dump
|
||||
(0.6 s here) and a load that must not trip on the two objects the new engine's entrypoint already made (the
|
||||
bootstrap role and database: measured, `pg_dumpall` loaded raw gives exactly two `already exists` ERRORs).
|
||||
**Why (b):** decision 16 says "save everything". The two known collisions are removed precisely — the
|
||||
entrypoint's empty databases are dropped before the load (only when they hold no table), and the dump's
|
||||
`CREATE ROLE <name>;` line is skipped for a role that already exists (its `ALTER ROLE … PASSWORD` still
|
||||
runs) — so any OTHER error stops the load and the undo runs. Proven by a test and live.
|
||||
|
||||
## A3 — the check
|
||||
|
||||
Old engine, before it stops; new engine, after the load — the same query set (0.5 s each here):
|
||||
per database: owner, encoding, collation; every role (superuser, login, has a password); every extension by
|
||||
name; every table's exact row count (`count(*)` per table). After: plus `PG_VERSION` of the new datadir
|
||||
(`$PGDATA/PG_VERSION`, which follows 17's and 18's different layouts). All IDENTICAL for docmost on both
|
||||
load routes.
|
||||
|
||||
## A4 — the undo after the volume was EMPTIED: proven by hand
|
||||
|
||||
`A4-undo-after-empty.txt`: copy with the product's own helper command (+ marker) → volume emptied (0
|
||||
entries) → the product's `Restore` command → `PG_VERSION` 16 → the old engine "ready to accept
|
||||
connections" → **the seeded account logged in through docmost's front door** (`A4-after-undo-productstart`).
|
||||
**The brief's claim is TRUE.**
|
||||
|
||||
*Recorded as it happened:* my first start after the restore was a bare `docker compose up -d` in the stack
|
||||
dir, which has no `.env` (the controller passes the app's env itself). docmost started without its secrets,
|
||||
crash-looped, and **the box stopped it after 7 restarts in 10 min (decision 28) — the product working**, not
|
||||
a fault. Started again with the product's own Start (`POST /api/stacks/docmost/start`); the seed read back.
|
||||
|
||||
## A5 — space
|
||||
|
||||
The conversion adds, on top of today's undo copy (the old datadir, 68.7 MB here): the dump (132 KB — 0.2 % of
|
||||
the datadir) and nothing for the new datadir (it is built IN the emptied volume: 52.3 MB). The box's check:
|
||||
**the dump's bound is the DB volume's own size** (a logical dump of live rows is smaller than the datadir
|
||||
holding them, measured 0.2 %), with a 25 % margin, on the filesystem that holds the stack directory (where the
|
||||
dump is written), plus the fixed 2 GB floor the update already keeps. Refused before anything moves.
|
||||
@@ -0,0 +1,55 @@
|
||||
### RED-PROOF B1 no mark → refused — mutation in pgconvert.go
|
||||
--- FAIL: TestConvert_MajorWithoutMarkIsRefused (0.00s)
|
||||
pgconvert_test.go:215: the preflight let it through: <nil>
|
||||
--- FAIL: TestConvert_JobRefusesAMarkThatDoesNotMatch (0.01s)
|
||||
pgconvert_test.go:235: phase "done" key "", want failed with the no-test sentence
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.020s
|
||||
### RED-PROOF B2 marker validated again before emptying — mutation in pgconvert.go
|
||||
--- FAIL: TestConvert_NeverEmptiedWithoutMarker (0.01s)
|
||||
pgconvert_test.go:338: Empty was called although the copy had no marker
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.015s
|
||||
FAIL
|
||||
### RED-PROOF B3 counts compared — mutation in pgconvert.go
|
||||
--- FAIL: TestConvert_CountsDifferAreUndone (0.01s)
|
||||
pgconvert_test.go:271: ended "done" (), want undone
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.021s
|
||||
FAIL
|
||||
### RED-PROOF B3 PG_VERSION checked — mutation in pgconvert.go
|
||||
--- FAIL: TestConvert_WrongVersionIsUndone (0.01s)
|
||||
pgconvert_test.go:304: ended "done" (), want undone
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.019s
|
||||
FAIL
|
||||
### RED-PROOF B3 cut-off dump refused — mutation in pgconvert.go
|
||||
--- FAIL: TestConvert_CutOffDumpNeverEmptiesTheVolume (0.01s)
|
||||
pgconvert_test.go:315: ended "done" (), want undone
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.018s
|
||||
FAIL
|
||||
### RED-PROOF B4 restart during converting → undo — mutation in update.go
|
||||
--- FAIL: TestConvert_RestartDuringConvertingIsUndone (0.06s)
|
||||
pgconvert_test.go:399: RecoverUpdates resumed []
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.062s
|
||||
FAIL
|
||||
### RED-PROOF B5 old datadir copy kept — mutation in update.go
|
||||
--- FAIL: TestConvert_HappyPath (0.01s)
|
||||
pgconvert_test.go:196: the kept copy is not recorded: <nil>
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s
|
||||
FAIL
|
||||
### RED-PROOF B5 release wired — mutation in main.go
|
||||
--- FAIL: TestConvert_ReleaseIsWiredAtStartup (0.01s)
|
||||
conversion_wiring_test.go:11: ReleaseConversionCopies is never called — a converted app's old datadir copy would stay on disk for ever
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.015s
|
||||
FAIL
|
||||
### RED-PROOF A5 space checked before anything moves — mutation in update.go
|
||||
--- FAIL: TestConvert_NoSpaceIsRefusedWithNothingMoved (0.01s)
|
||||
pgconvert_test.go:353: phase "done" key ""
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s
|
||||
FAIL
|
||||
@@ -0,0 +1,18 @@
|
||||
--- PASS: TestConvert_HappyPath (0.01s)
|
||||
--- PASS: TestConvert_MajorWithoutMarkIsRefused (0.01s)
|
||||
--- PASS: TestConvert_JobRefusesAMarkThatDoesNotMatch (0.00s)
|
||||
--- PASS: TestConvert_CountsDifferAreUndone (0.01s)
|
||||
--- PASS: TestConvert_LoadFailureIsUndone (0.01s)
|
||||
--- PASS: TestConvert_HealthFailureIsUndone (0.01s)
|
||||
--- PASS: TestConvert_WrongVersionIsUndone (0.01s)
|
||||
--- PASS: TestConvert_CutOffDumpNeverEmptiesTheVolume (0.01s)
|
||||
--- PASS: TestConvert_NeverEmptiedWithoutMarker (0.01s)
|
||||
--- PASS: TestConvert_NoSpaceIsRefusedWithNothingMoved (0.00s)
|
||||
--- PASS: TestConvert_RestartDuringConvertingIsUndone (0.06s)
|
||||
--- PASS: TestConvert_ReleaseAfterAProvenBackup (0.01s)
|
||||
--- PASS: TestConvert_PostgresMajorFromTags (0.00s)
|
||||
--- PASS: TestConvert_PostgresFamilyMatchesTheBackupSide (0.00s)
|
||||
--- PASS: TestConvert_FilterSkipsOnlyExactLines (0.00s)
|
||||
--- PASS: TestConvert_ValidateDumpAll (0.00s)
|
||||
--- PASS: TestConvert_MarkDoesNotChangeOldLadderPrints (0.00s)
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.189s
|
||||
@@ -0,0 +1,13 @@
|
||||
local dir active 40453376 22728440 15637820 56.18%
|
||||
nvme-scratch dir active 983379700 75082704 858270384 7.64%
|
||||
template debian-13-standard_13.6-1_arm64.tar.zst
|
||||
sync_wait: 34 An error occurred in another process (expected sequence number 7)
|
||||
__lxc_start: 2288 Failed to spawn container "9401"
|
||||
startup for container '9401' failed
|
||||
purging CT 9401 from related configurations..
|
||||
destroyed wrong-arch 9401
|
||||
debian-13-standard_13.6-1_amd64.tar.zst
|
||||
192.168.0.121/24
|
||||
Docker version 26.1.5+dfsg1, build a72d7cd
|
||||
Docker Compose version 2.26.1-4
|
||||
net-ok
|
||||
@@ -0,0 +1,11 @@
|
||||
ok 18 moves the data mount to /var/lib/postgresql
|
||||
ok 18 leaves the redis mount alone
|
||||
ok 16 keeps /var/lib/postgresql/data
|
||||
ok back from 18 to 16 restores /data
|
||||
ok pg_conversion names the one moving service
|
||||
ok within a major is no conversion
|
||||
ok pgvector's pg16 tag reads 16
|
||||
ok an entry with a good mark is well-formed
|
||||
ok a mark naming a service outside `to` is refused
|
||||
ok a mark going DOWN is refused
|
||||
pg conversion tests: 0 failed
|
||||
@@ -0,0 +1,8 @@
|
||||
ok FACT: docmost-postgres 16-alpine -> 17-alpine bundled with the app bump rc=1 (expected 1)
|
||||
ok FACT: docmost-postgres postgres:16-alpine -> 17-alpine rc=1 (expected 1)
|
||||
ok GENUINE: docmost-postgres 16 -> 18 ALONE with its proven two-venue MARKED entry rc=0 (expected 0)
|
||||
ok DECOY: the entry is proven on both venues but carries NO conversion mark rc=1 (expected 1)
|
||||
ok DECOY: the marked entry cites ONE venue only (no box_evidence) rc=1 (expected 1)
|
||||
ok DECOY: the mark sits in a COMMENT, the entry has none rc=1 (expected 1)
|
||||
ok DECOY: the marked entry, but the move is BUNDLED with the app's own bump rc=1 (expected 1)
|
||||
ok FACT: adventurelog's postgis 16 -> 17 (the postgis family was never judged before) rc=1 (expected 1)
|
||||
@@ -0,0 +1,7 @@
|
||||
# RED-PROOF of check-engine-major.py's PostgreSQL clause (2026-09-25, catalog tree before commit)
|
||||
# mutation: `if why is None and not others:` -> `if not others:` (the ladder proof is not consulted)
|
||||
# python3 scripts/test_gate_decoys.py, lines quoted verbatim:
|
||||
FAIL: FACT: docmost-postgres 16-alpine -> 17-alpine bundled with the app bump: rc=0 expected 1; missing ['ENGINE-MAJOR GATE FAILED']
|
||||
FAIL: FACT: docmost-postgres postgres:16-alpine -> 17-alpine: rc=0 expected 1; missing ['R-463']
|
||||
FAIL: DECOY: the entry is proven on both venues but carries NO conversion mark: rc=0 expected 1; missing ['NOT PROVEN FOR THIS APP', 'engine_conversion None']
|
||||
# tree restored from the saved copy; the suite green again: C2-decoys-green.txt
|
||||
@@ -0,0 +1,5 @@
|
||||
# 9202 before: gitea.dooplex.hu/admin/felhom-controller:0.272.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
|
||||
# 9202 after: gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 healthy
|
||||
2026/09/25 11:04:49 main.go:336: [INFO] felhom-controller 0.273.0-rc1 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/09/25 11:04:49 undo.go:675: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
|
||||
@@ -0,0 +1,10 @@
|
||||
git:
|
||||
branch: main
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: "admin"
|
||||
hub:
|
||||
update:
|
||||
health_timeout: 90s
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
# D1-happy-path — 2026-09-25T13:12:42+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
|
||||
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
before: seed reads back through the front door = True
|
||||
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
|
||||
phase +2.1s pulling | Új verzió letöltése… | err=None
|
||||
phase +3.2s copying | Az adatok másolása a frissítés előtt… | err=None
|
||||
phase +5.8s converting | Adatbázis átalakítása | err=None
|
||||
phase +15.8s starting | Indítás az új verzióval… | err=None
|
||||
phase +26.8s verifying | Működés ellenőrzése… | err=None
|
||||
phase +42.1s done | Frissítve | err=None
|
||||
result: {"accepted": true, "http": "202", "duration_s": 42.1, "final_phase": "done", "update_error": null, "hold_reason": null, "state": "running"}
|
||||
after: PG_VERSION/image=['18', 'postgres:18-alpine@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy={'volume': 'docmost_docmost_postgres_data', 'copy': 'docmost_docmost_postgres_data.pre-update-20260925T111256Z', 'at': '2026-09-25T11:13:34Z', 'from': 16, 'to': 18}
|
||||
after: seed reads back through the front door = True
|
||||
controller lines (docmost / conversion):
|
||||
2026/09/25 11:12:52 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
|
||||
2026/09/25 11:12:52 update.go:1365: [INFO] [stacks] update docmost: phase checking
|
||||
2026/09/25 11:12:52 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
|
||||
2026/09/25 11:12:52 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
|
||||
2026/09/25 11:12:52 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:11:16Z (2m0s old, limit 24h0m0s)
|
||||
2026/09/25 11:12:52 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
|
||||
2026/09/25 11:12:53 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T111253Z-docmost-postgres.sql]
|
||||
2026/09/25 11:12:54 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.4 MiB
|
||||
2026/09/25 11:12:54 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
|
||||
2026/09/25 11:12:54 update.go:1365: [INFO] [stacks] update docmost: phase pinning
|
||||
2026/09/25 11:12:54 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
|
||||
2026/09/25 11:12:54 update.go:1365: [INFO] [stacks] update docmost: phase pulling
|
||||
2026/09/25 11:12:55 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:12:56 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:12:57 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111256Z in 829ms
|
||||
2026/09/25 11:12:57 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:12:58 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111256Z in 440ms
|
||||
2026/09/25 11:12:58 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:12:58 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111256Z in 418ms
|
||||
2026/09/25 11:12:58 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:12:58 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111256Z)
|
||||
2026/09/25 11:13:02 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 566ms — 131 kB, completion line present; before: 2 database(s), 48 table(s), 71 row(s)
|
||||
2026/09/25 11:13:02 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:13:03 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111256Z holds the old datadir, marker checked twice)
|
||||
2026/09/25 11:13:07 pgconvert.go:459: [INFO] [stacks] update docmost: loaded into 18 in 1.753s (dropped the entrypoint's empty database(s) [docmost]; skipped CREATE ROLE for [docmost])
|
||||
2026/09/25 11:13:08 pgconvert.go:475: [INFO] [stacks] update docmost: CONVERTED 16 → 18 in 9.942s — the check is equal (2 database(s), 48 table(s), 71 row(s)), PG_VERSION 18
|
||||
2026/09/25 11:13:08 update.go:1365: [INFO] [stacks] update docmost: phase starting
|
||||
2026/09/25 11:13:19 update.go:1365: [INFO] [stacks] update docmost: phase verifying
|
||||
2026/09/25 11:13:34 update.go:997: [INFO] [stacks] update docmost: healthy after 15s (the app's health check passed)
|
||||
2026/09/25 11:13:34 update.go:1015: [INFO] [stacks] update docmost: KEEPING the pre-conversion datadir copy docmost_docmost_postgres_data.pre-update-20260925T111256Z (PostgreSQL 16) until a backup of the converted app is proven
|
||||
2026/09/25 11:13:34 update.go:1033: [INFO] [stacks] update docmost: DONE in 42s
|
||||
|
||||
undo copies now: docmost_docmost_postgres_data.pre-update-20260925T111256Z
|
||||
done
|
||||
@@ -0,0 +1,9 @@
|
||||
# 'Mentés most' on docmost after the conversion (the release of the kept copy waits for a PROVEN backup) 2026-09-25T13:13:59+0200
|
||||
13:13:59 [4] backup press SKIPPED for docmost (R-648: whole-box only; the update's backing-up phase backs up docmost alone)
|
||||
None
|
||||
-rw-r--r-- 1 root root 154767 2026-09-25T11:06:21 pre-restore-20260925T110621Z-docmost-postgres.sql
|
||||
-rw-r--r-- 1 root root 154903 2026-09-25T11:08:02 pre-restore-20260925T110802Z-docmost-postgres.sql
|
||||
-rw-r--r-- 1 root root 155357 2026-09-25T11:11:16 pre-restore-20260925T111115Z-docmost-postgres.sql
|
||||
-rw-r--r-- 1 root root 155810 2026-09-25T11:12:53 pre-restore-20260925T111253Z-docmost-postgres.sql
|
||||
docmost_docmost_postgres_data.pre-update-20260925T111256Z
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
# Da-load-fails — 2026-09-25T13:06:11+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
|
||||
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
before: seed reads back through the front door = True
|
||||
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
|
||||
phase +2.1s pulling | Új verzió letöltése… | err=None
|
||||
phase +2.7s copying | Az adatok másolása a frissítés előtt… | err=None
|
||||
phase +5.8s converting | Adatbázis átalakítása | err=None
|
||||
phase +13.7s undoing | Visszaállítás az előző változatra… | err=None
|
||||
phase +43.6s undone | Visszaállítva az előző változatra | err=None
|
||||
result: {"accepted": true, "http": "202", "duration_s": 43.6, "final_phase": "undone", "update_error": null, "hold_reason": null, "state": "running"}
|
||||
after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None
|
||||
after: seed reads back through the front door = True
|
||||
controller lines (docmost / conversion):
|
||||
2026/09/25 11:06:21 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
|
||||
2026/09/25 11:06:21 update.go:1365: [INFO] [stacks] update docmost: phase checking
|
||||
2026/09/25 11:06:21 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
|
||||
2026/09/25 11:06:21 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
|
||||
2026/09/25 11:06:21 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:04:49Z (2m0s old, limit 24h0m0s)
|
||||
2026/09/25 11:06:21 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
|
||||
2026/09/25 11:06:21 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T110621Z-docmost-postgres.sql]
|
||||
2026/09/25 11:06:22 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.3 MiB
|
||||
2026/09/25 11:06:23 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
|
||||
2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase pinning
|
||||
2026/09/25 11:06:23 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
|
||||
2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase pulling
|
||||
2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:06:24 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:06:25 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T110624Z in 869ms
|
||||
2026/09/25 11:06:25 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:06:26 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T110624Z in 431ms
|
||||
2026/09/25 11:06:26 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:06:26 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T110624Z in 422ms
|
||||
2026/09/25 11:06:26 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:06:26 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T110624Z)
|
||||
2026/09/25 11:06:30 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 572ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 65 row(s)
|
||||
2026/09/25 11:06:30 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:06:31 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T110624Z holds the old datadir, marker checked twice)
|
||||
2026/09/25 11:06:34 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: conversion failed: loading the dump into 18: exit status 3 (psql: ERROR: extension "adminpack" is not available)
|
||||
2026/09/25 11:06:34 update.go:1058: [INFO] [stacks] update docmost: kept 11240 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T110634Z/compose-logs.txt before stopping it (R-621)
|
||||
2026/09/25 11:06:34 update.go:1365: [INFO] [stacks] update docmost: phase undoing
|
||||
2026/09/25 11:06:34 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: conversion failed: loading the dump into 18: exit status 3 (psql: ERROR: extension "adminpack" is not available))
|
||||
2026/09/25 11:07:04 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
|
||||
2026/09/25 11:07:04 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone
|
||||
2026/09/25 11:07:04 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 30s — the previous version is running on the data from before the update (the app's health check passed)
|
||||
|
||||
undo copies now:
|
||||
done
|
||||
@@ -0,0 +1,6 @@
|
||||
# case (a) setup: an object 16 has and 18 does not — the adminpack extension (removed in PostgreSQL 17)
|
||||
CREATE EXTENSION
|
||||
adminpack pg_trgm plpgsql unaccent
|
||||
control: 18 has no adminpack: 0
|
||||
# after case (a): adminpack removed again
|
||||
DROP EXTENSION
|
||||
@@ -0,0 +1,53 @@
|
||||
# Db-health-fails — 2026-09-25T13:07:52+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
|
||||
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
before: seed reads back through the front door = True
|
||||
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
|
||||
phase +2.1s pulling | Új verzió letöltése… | err=None
|
||||
phase +2.6s copying | Az adatok másolása a frissítés előtt… | err=None
|
||||
phase +5.8s converting | Adatbázis átalakítása | err=None
|
||||
phase +15.2s starting | Indítás az új verzióval… | err=None
|
||||
phase +26.3s verifying | Működés ellenőrzése… | err=None
|
||||
phase +117.1s undoing | Visszaállítás az előző változatra… | err=None
|
||||
phase +148.1s undone | Visszaállítva az előző változatra | err=None
|
||||
result: {"accepted": true, "http": "202", "duration_s": 148.1, "final_phase": "undone", "update_error": null, "hold_reason": null, "state": "running"}
|
||||
after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None
|
||||
after: seed reads back through the front door = True
|
||||
controller lines (docmost / conversion):
|
||||
2026/09/25 11:08:02 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
|
||||
2026/09/25 11:08:02 update.go:1365: [INFO] [stacks] update docmost: phase checking
|
||||
2026/09/25 11:08:02 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
|
||||
2026/09/25 11:08:02 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
|
||||
2026/09/25 11:08:02 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:06:21Z (2m0s old, limit 24h0m0s)
|
||||
2026/09/25 11:08:02 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
|
||||
2026/09/25 11:08:02 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T110802Z-docmost-postgres.sql]
|
||||
2026/09/25 11:08:03 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.3 MiB
|
||||
2026/09/25 11:08:04 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
|
||||
2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase pinning
|
||||
2026/09/25 11:08:04 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
|
||||
2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase pulling
|
||||
2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:08:05 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:08:06 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T110805Z in 839ms
|
||||
2026/09/25 11:08:06 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:08:07 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T110805Z in 435ms
|
||||
2026/09/25 11:08:07 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:08:07 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T110805Z in 441ms
|
||||
2026/09/25 11:08:07 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:08:07 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T110805Z)
|
||||
2026/09/25 11:08:10 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 566ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 67 row(s)
|
||||
2026/09/25 11:08:11 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:08:12 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T110805Z holds the old datadir, marker checked twice)
|
||||
2026/09/25 11:08:16 pgconvert.go:459: [INFO] [stacks] update docmost: loaded into 18 in 1.747s (dropped the entrypoint's empty database(s) [docmost]; skipped CREATE ROLE for [docmost])
|
||||
2026/09/25 11:08:17 pgconvert.go:475: [INFO] [stacks] update docmost: CONVERTED 16 → 18 in 9.866s — the check is equal (2 database(s), 48 table(s), 67 row(s)), PG_VERSION 18
|
||||
2026/09/25 11:08:17 update.go:1365: [INFO] [stacks] update docmost: phase starting
|
||||
2026/09/25 11:08:28 update.go:1365: [INFO] [stacks] update docmost: phase verifying
|
||||
2026/09/25 11:09:58 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: not healthy: not healthy within 1m30s (last: health check failing)
|
||||
2026/09/25 11:09:59 update.go:1058: [INFO] [stacks] update docmost: kept 13077 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T110958Z/compose-logs.txt before stopping it (R-621)
|
||||
2026/09/25 11:09:59 update.go:1365: [INFO] [stacks] update docmost: phase undoing
|
||||
2026/09/25 11:09:59 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))
|
||||
2026/09/25 11:10:30 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
|
||||
2026/09/25 11:10:30 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone
|
||||
2026/09/25 11:10:30 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 31s — the previous version is running on the data from before the update (the app's health check passed)
|
||||
|
||||
undo copies now:
|
||||
done
|
||||
@@ -0,0 +1,83 @@
|
||||
# Dc-restart-mid-convert — 2026-09-25T13:11:04+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
|
||||
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
before: seed reads back through the front door = True
|
||||
[kill] 2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice) | KILLED controller pid 554023 at 11:11:24.114456007
|
||||
[kill] waiting for the controller to come back and the journal to be resumed
|
||||
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
|
||||
phase +2.1s pulling | Új verzió letöltése… | err=None
|
||||
phase +2.6s copying | Az adatok másolása a frissítés előtt… | err=None
|
||||
phase +5.8s converting | Adatbázis átalakítása | err=None
|
||||
phase +8.4s None | None | err=None
|
||||
result: {"accepted": true, "http": "202", "duration_s": 8.5, "final_phase": null, "update_error": null, "hold_reason": null, "state": null, "after_restart": {"phase": "undone", "error": null, "hold": null}}
|
||||
after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None
|
||||
after: seed reads back through the front door = True
|
||||
controller lines (docmost / conversion):
|
||||
2026/09/25 11:11:15 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
|
||||
2026/09/25 11:11:15 update.go:1365: [INFO] [stacks] update docmost: phase checking
|
||||
2026/09/25 11:11:15 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
|
||||
2026/09/25 11:11:15 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
|
||||
2026/09/25 11:11:15 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:08:02Z (3m0s old, limit 24h0m0s)
|
||||
2026/09/25 11:11:15 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
|
||||
2026/09/25 11:11:16 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T111115Z-docmost-postgres.sql]
|
||||
2026/09/25 11:11:17 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.4 MiB
|
||||
2026/09/25 11:11:17 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
|
||||
2026/09/25 11:11:17 update.go:1365: [INFO] [stacks] update docmost: phase pinning
|
||||
2026/09/25 11:11:17 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
|
||||
2026/09/25 11:11:17 update.go:1365: [INFO] [stacks] update docmost: phase pulling
|
||||
2026/09/25 11:11:18 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:11:19 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111119Z in 794ms
|
||||
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111119Z in 415ms
|
||||
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:11:21 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111119Z in 423ms
|
||||
2026/09/25 11:11:21 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:11:21 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111119Z)
|
||||
2026/09/25 11:11:22 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 539ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 69 row(s)
|
||||
2026/09/25 11:11:23 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice)
|
||||
2026/09/25 11:11:24 update.go:1432: [WARN] [stacks] update recovery: docmost was interrupted while UNDOING (started 2026-09-25T11:11:15Z) — marking it Updating and RESUMING the undo
|
||||
2026/09/25 11:11:24 main.go:496: [WARN] [update] 1 interrupted update(s) will resume after the backup side is wired: [docmost]
|
||||
2026/09/25 11:11:24 main.go:596: [WARN] [update] resumed 1 interrupted update(s)
|
||||
2026/09/25 11:11:24 update.go:1489: [INFO] [stacks] update docmost: resuming the UNDO after a controller restart (was converting)
|
||||
2026/09/25 11:11:24 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: the controller restarted during the database conversion — undoing it
|
||||
2026/09/25 11:11:25 update.go:1058: [INFO] [stacks] update docmost: kept 5297 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T111124Z/compose-logs.txt before stopping it (R-621)
|
||||
2026/09/25 11:11:25 update.go:1365: [INFO] [stacks] update docmost: phase undoing
|
||||
2026/09/25 11:11:25 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: the controller restarted during the database conversion — undoing it)
|
||||
2026/09/25 11:12:05 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
|
||||
2026/09/25 11:12:05 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone
|
||||
2026/09/25 11:12:05 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 40s — the previous version is running on the data from before the update (the app's health check passed)
|
||||
|
||||
undo copies now:
|
||||
done
|
||||
# the whole restart, from the controller's own log
|
||||
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111119Z in 794ms
|
||||
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111119Z in 415ms
|
||||
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/25 11:11:21 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111119Z in 423ms
|
||||
2026/09/25 11:11:21 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:11:21 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111119Z)
|
||||
2026/09/25 11:11:22 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 539ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 69 row(s)
|
||||
2026/09/25 11:11:23 update.go:1365: [INFO] [stacks] update docmost: phase converting
|
||||
2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice)
|
||||
2026/09/25 11:11:24 main.go:336: [INFO] felhom-controller 0.273.0-rc1 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/09/25 11:11:24 update.go:1432: [WARN] [stacks] update recovery: docmost was interrupted while UNDOING (started 2026-09-25T11:11:15Z) — marking it Updating and RESUMING the undo
|
||||
2026/09/25 11:11:24 main.go:496: [WARN] [update] 1 interrupted update(s) will resume after the backup side is wired: [docmost]
|
||||
2026/09/25 11:11:24 main.go:596: [WARN] [update] resumed 1 interrupted update(s)
|
||||
2026/09/25 11:11:24 update.go:1489: [INFO] [stacks] update docmost: resuming the UNDO after a controller restart (was converting)
|
||||
2026/09/25 11:11:24 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: the controller restarted during the database conversion — undoing it
|
||||
2026/09/25 11:11:25 update.go:1058: [INFO] [stacks] update docmost: kept 5297 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T111124Z/compose-logs.txt before stopping it (R-621)
|
||||
2026/09/25 11:11:25 update.go:1365: [INFO] [stacks] update docmost: phase undoing
|
||||
2026/09/25 11:11:25 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: the controller restarted during the database conversion — undoing it)
|
||||
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
|
||||
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
|
||||
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
|
||||
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
|
||||
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
|
||||
2026/09/25 11:11:34 healthprobe.go:190: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host
|
||||
2026/09/25 11:11:38 [INFO] [stacks] SaveAppConfig: saved config for docmost
|
||||
2026/09/25 11:11:38 pin.go:93: [INFO] [stacks] pin docmost: docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:16-alpine, docmost-redis=redis:7-alpine
|
||||
2026/09/25 11:11:44 healthprobe.go:190: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host
|
||||
2026/09/25 11:12:05 [INFO] [stacks] SaveAppConfig: saved config for docmost
|
||||
2026/09/25 11:12:05 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
|
||||
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"app": "docmost",
|
||||
"venue": "box 9202 (drill catalog, controller 0.273.0-rc1), the product's guarded Update",
|
||||
"from": {
|
||||
"docmost": "docmost/docmost:0.96.0",
|
||||
"docmost-postgres": "postgres:16-alpine",
|
||||
"docmost-redis": "redis:7-alpine"
|
||||
},
|
||||
"to": {
|
||||
"docmost": "docmost/docmost:0.96.0",
|
||||
"docmost-postgres": "postgres:18-alpine",
|
||||
"docmost-redis": "redis:7-alpine"
|
||||
},
|
||||
"verdict": "proven",
|
||||
"seed_read_before": true,
|
||||
"seed_read_after": true,
|
||||
"healthy_after": true,
|
||||
"engine_conversion": {
|
||||
"service": "docmost-postgres",
|
||||
"engine": "postgres",
|
||||
"from": 16,
|
||||
"to": 18,
|
||||
"result": "converted",
|
||||
"controller_line": "CONVERTED 16 → 18 in 9.942s — the check is equal (2 database(s), 48 table(s), 71 row(s)), PG_VERSION 18"
|
||||
},
|
||||
"duration_s": 42.1,
|
||||
"measured_at": "2026-09-25T11:13:34Z",
|
||||
"evidence": "felhom.eu/documentation/audits/night-2026-09-26/D/D1-happy-path.txt"
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
12:40:28 before:
|
||||
2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-26 05:30 CEST
|
||||
2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-26 04:00 CEST
|
||||
2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-26 03:30 CEST
|
||||
12:40:28 POST /backups/window window_start=12:44 -> 303
|
||||
12:40:34 after:
|
||||
2026/09/25 10:40:28 auth.go:142: [DEBUG] [web] auth: valid session for POST /backups/window
|
||||
2026/09/25 10:40:28 server.go:545: [DEBUG] [web] ServeHTTP: POST /backups/window from 172.18.0.6:60864
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST)
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 13:44 (next run 2026-09-25 13:44 CEST)
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 14:29 (next run 2026-09-25 14:29 CEST)
|
||||
2026/09/25 10:40:28 backup_handlers.go:65: [INFO] [web] backup window set to 12:44 (legs 12:44/13:44/14:29)
|
||||
@@ -0,0 +1,68 @@
|
||||
Fri Sep 25 10:40:56 UTC 2026
|
||||
2026/09/25 10:38:19 router.go:946: [ERROR] [api] Remove failed for nextcloud: A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
|
||||
2026/09/25 10:38:19 router.go:907: [INFO] [api] Remove requested for stack: nextcloud
|
||||
2026/09/25 10:38:19 delete.go:597: [INFO] Removing deployed stack: nextcloud (removeHDDData=false, hddDeclared=true, backupPaths=2)
|
||||
2026/09/25 10:38:19 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/25 10:38:47 delete.go:650: [INFO] [stacks] RemoveStack nextcloud: removed volume(s) [nextcloud_nextcloud_db_data nextcloud_nextcloud_html nextcloud_nextcloud_redis_data]
|
||||
2026/09/25 10:38:47 delete.go:749: [INFO] Stack nextcloud removed successfully (took 27.4s)
|
||||
2026/09/25 10:39:01 router.go:431: [INFO] [api] Deploy requested for stack: nextcloud
|
||||
2026/09/25 10:39:01 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
|
||||
2026/09/25 10:39:01 deploy.go:391: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH]
|
||||
2026/09/25 10:39:01 manager.go:1592: [INFO] [stacks] Deploying stack nextcloud — checking 3 images...
|
||||
2026/09/25 10:39:13 deploy.go:481: [INFO] [stacks] Stack nextcloud deployed successfully (took 11.2s)
|
||||
2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
|
||||
2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
|
||||
2026/09/25 10:39:13 installed.go:420: [INFO] [stacks] installed-images nextcloud: recorded 3 service(s) (nextcloud=nextcloud:34.0.4-apache (sha256:a5ace30c695a…), nextcloud-db=mariadb:12.3 (sha256:805c8e104bd5…), nextcloud-red
|
||||
2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
|
||||
2026/09/25 10:39:13 pin.go:93: [INFO] [stacks] pin nextcloud: nextcloud=nextcloud:34.0.4-apache, nextcloud-db=mariadb:12.3, nextcloud-redis=redis:7-alpine
|
||||
2026/09/25 10:39:16 manager.go:1553: [INFO] [stacks] Stack nextcloud post-start status:
|
||||
2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
|
||||
2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 14 seconds (healthy)
|
||||
2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 14 seconds (healthy)
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST)
|
||||
66548367016e nextcloud-db nextcloud mariadb:12.3
|
||||
5435a635557b nextcloud-redis nextcloud redis:7-alpine
|
||||
2026/09/25 10:40:28 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
|
||||
2026/09/25 10:40:28 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withhe
|
||||
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud:
|
||||
total 16
|
||||
drwxr-xr-x 3 root root 4096 10:40:28 .
|
||||
drwxr-xr-x 3 root root 4096 10:40:28 ..
|
||||
drwxr-xr-x 2 root root 4096 10:40:28 compose
|
||||
-rw-r--r-- 1 root root 1439 10:40:28 manifest.json
|
||||
## after the db-dump leg
|
||||
2026/09/25 10:44:23 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_db_data → 164.5 MB
|
||||
2026/09/25 10:44:28 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 761.6 MB
|
||||
2026/09/25 10:44:28 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 187.5 KB
|
||||
2026/09/25 10:44:28 backup.go:933: [INFO] [backup] Restarting nextcloud after volume dump
|
||||
2026/09/25 10:44:28 manager.go:1207: [INFO] [stacks] Starting stack: nextcloud
|
||||
2026/09/25 10:44:35 manager.go:1222: [INFO] [stacks] Stack nextcloud started successfully (took 6.2s)
|
||||
2026/09/25 10:44:38 manager.go:1553: [INFO] [stacks] Stack nextcloud post-start status:
|
||||
2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (healt
|
||||
2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 9 seconds (healt
|
||||
2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 9 seconds (healt
|
||||
2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0
|
||||
2026/09/25 10:44:56 scheduler.go:363: [INFO] [scheduler] Job db-dump completed (took 56.675s)
|
||||
-rw-r--r-- 1 root root 1587 10:44:56 /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/manifest.json
|
||||
|
||||
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/compose:
|
||||
total 32
|
||||
drwxr-xr-x 2 root root 4096 10:44:56 .
|
||||
drwxr-xr-x 5 root root 4096 10:44:56 ..
|
||||
-rw-r--r-- 1 root root 9180 10:44:56 .felhom.yml
|
||||
-rw------- 1 root root 611 10:44:56 app.yaml
|
||||
-rw-r--r-- 1 root root 4809 10:44:56 docker-compose.yml
|
||||
|
||||
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps:
|
||||
total 864
|
||||
drwxr-xr-x 2 root root 4096 10:44:00 .
|
||||
drwxr-xr-x 5 root root 4096 10:44:56 ..
|
||||
-rw-r--r-- 1 root root 873120 10:44:00 nextcloud-mariadb.sql
|
||||
|
||||
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/volume-dumps:
|
||||
total 948440
|
||||
drwxr-xr-x 2 root root 4096 10:44:28 .
|
||||
drwxr-xr-x 5 root root 4096 10:44:56 ..
|
||||
-rw-r--r-- 1 root root 172453376 10:44:22 nextcloud_nextcloud_db_data.tar
|
||||
-rw-r--r-- 1 root root 798543360 10:44:26 nextcloud_nextcloud_html.tar
|
||||
-rw-r--r-- 1 root root 192000 10:44:28 nextcloud_nextcloud_redis_data.tar
|
||||
@@ -0,0 +1,12 @@
|
||||
12:45:23 before:
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST)
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 13:44 (next run 2026-09-25 13:44 CEST)
|
||||
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 14:29 (next run 2026-09-25 14:29 CEST)
|
||||
12:45:23 POST /backups/window window_start=02:30 -> 303
|
||||
12:45:29 after:
|
||||
2026/09/25 10:45:23 auth.go:142: [DEBUG] [web] auth: valid session for POST /backups/window
|
||||
2026/09/25 10:45:23 server.go:545: [DEBUG] [web] ServeHTTP: POST /backups/window from 172.18.0.6:39110
|
||||
2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 12:44 → 02:30 (next run 2026-09-26 02:30 CEST)
|
||||
2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 13:44 → 03:30 (next run 2026-09-26 03:30 CEST)
|
||||
2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 14:29 → 04:15 (next run 2026-09-26 04:15 CEST)
|
||||
2026/09/25 10:45:23 backup_handlers.go:65: [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15)
|
||||
@@ -0,0 +1,37 @@
|
||||
('200', {'ok': True, 'message': 'Stack nextcloud stop completed'})
|
||||
remove keep data + keep backups -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)'], 'verified': True}, 'message': 'Stack nextcloud removed'}
|
||||
total 28
|
||||
drwxrwx--- 5 www-data www-data 4096 Sep 25 10:39 .
|
||||
drwxr-xr-x 4 root root 4096 Sep 25 10:39 ..
|
||||
-rw-rw-r-- 1 www-data www-data 542 Sep 25 10:39 .htaccess
|
||||
-rw-rw-r-- 1 www-data www-data 52 Sep 25 10:39 .ncdata
|
||||
drwxr-xr-x 3 www-data www-data 4096 Sep 25 10:39 admin
|
||||
drwxr-xr-x 4 www-data www-data 4096 Sep 25 10:39 appdata_ocrpy9goog25
|
||||
drwxr-xr-x 4 www-data www-data 4096 Sep 25 10:40 drilld112bd
|
||||
-rw-rw-r-- 1 www-data www-data 0 Sep 25 10:39 index.html
|
||||
-rw-rw-r-- 1 www-data www-data 0 Sep 25 10:39 nextcloud.log
|
||||
/mnt/felhom-drives/scratch_hdd/appdata/nextcloud/admin/files:
|
||||
Documents
|
||||
Nextcloud Manual.pdf
|
||||
Nextcloud intro.mp4
|
||||
Nextcloud.png
|
||||
Photos
|
||||
Readme.md
|
||||
Reasons to use Nextcloud.pdf
|
||||
Templates
|
||||
Templates credits.md
|
||||
|
||||
/mnt/felhom-drives/scratch_hdd/appdata/nextcloud/drilld112bd/files:
|
||||
Documents
|
||||
Nextcloud Manual.pdf
|
||||
Nextcloud intro.mp4
|
||||
Nextcloud.png
|
||||
Photos
|
||||
Readme.md
|
||||
Reasons to use Nextcloud.pdf
|
||||
Templates
|
||||
Templates credits.md
|
||||
after-backup.txt
|
||||
before-backup.txt
|
||||
nextcloud
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# E1 Q1 — the removed-app row on /backups/restore (R-487)
|
||||
row present: True
|
||||
csolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → Biztonsági mentés — Visszaállítás enkisfelhom.hu Visszaállítás Alkalmazás: — Válasszon — Docmost Nextcloud Paperless-ngx <option value="pr
|
||||
snapshot_id in the row's form: []
|
||||
12:46:25 [R] restoring nextcloud from snapshot 'helyi' (of 0 offered)
|
||||
12:46:25 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started']
|
||||
12:46:25 + 0.0s restore (True, None, None)
|
||||
12:48:02 + 97.2s restore (False, None, None)
|
||||
12:48:02 [R] after restore: state=degraded hold=None phase=None
|
||||
{'http': 'HTTP/2 302', 'location': ['location: /backups/restore?flash=flash.restore.started'], 'seconds': 97.2, 'state_after': 'degraded', 'hold_after': None}
|
||||
12:58:16 app never answered on c-nc/status.php (last rc=0 code=404)
|
||||
12:58:16 nextcloud: the app never served /status.php
|
||||
12:58:16 E1-Q1-after-restore: read nextcloud=False
|
||||
12:58:16 PROPFIND files of drilld112bd -> 404; before-backup.txt visible=False; listing=[]
|
||||
12:58:16 PROPFIND files of drilld112bd -> 404; after-backup.txt visible=False; listing=[]
|
||||
state degraded
|
||||
Error response from daemon: container 96bf44773afc146b3fb41cf9f04b0faeb94a9b2a6fa02411af34a94da4ffa4f4 is not running
|
||||
/var/lib/docker/volumes/nextcloud_nextcloud_html/_data->/var/www/html /appdata/nextcloud->/var/www/html/data
|
||||
|
||||
## E1 Q1 evidence, 2026-09-25T10:58:39Z
|
||||
### env KEYS in the restored app.yaml (values never printed)
|
||||
|
||||
### env KEYS in the unit's app.yaml
|
||||
env DB_PASSWORD DOMAIN HDD_PATH MYSQL_ROOT_PASSWORD NEXTCLOUD_ADMIN_USER SUBDOMAIN
|
||||
### nextcloud container mounts
|
||||
/var/lib/docker/volumes/nextcloud_nextcloud_html/_data -> /var/www/html
|
||||
/appdata/nextcloud -> /var/www/html/data
|
||||
|
||||
total 12
|
||||
drwxr-xr-x 3 root root 4096 Sep 15 09:17 .
|
||||
drwxr-xr-x 20 root root 4096 Sep 24 20:52 ..
|
||||
### nextcloud-db log tail
|
||||
2026-09-25 12:56:23+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
|
||||
2026-09-25 12:56:23+02:00 [ERROR] [Entrypoint]: Database is uninitialized and password option is not specified
|
||||
You need to specify one of MARIADB_ROOT_PASSWORD, MARIADB_ROOT_PASSWORD_HASH, MARIADB_ALLOW_EMPTY_ROOT_PASSWORD and MARIADB_RANDOM_ROOT_PASSWORD
|
||||
2026-09-25 12:57:23+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
|
||||
2026-09-25 12:57:24+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB
|
||||
2026-09-25 12:57:24+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql'
|
||||
2026-09-25 12:57:24+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
|
||||
2026-09-25 12:57:24+02:00 [ERROR] [Entrypoint]: Database is uninitialized and password option is not specified
|
||||
You need to specify one of MARIADB_ROOT_PASSWORD, MARIADB_ROOT_PASSWORD_HASH, MARIADB_ALLOW_EMPTY_ROOT_PASSWORD and MARIADB_RANDOM_ROOT_PASSWORD
|
||||
2026-09-25 12:58:24+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
|
||||
2026-09-25 12:58:24+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB
|
||||
2026-09-25 12:58:24+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql'
|
||||
2026-09-25 12:58:24+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
|
||||
2026-09-25 12:58:25+02:00 [ERROR] [Entrypoint]: Database is uninitialized and password option is not specified
|
||||
You need to specify one of MARIADB_ROOT_PASSWORD, MARIADB_ROOT_PASSWORD_HASH, MARIADB_ALLOW_EMPTY_ROOT_PASSWORD and MARIADB_RANDOM_ROOT_PASSWORD
|
||||
### controller restore lines
|
||||
2026/09/25 10:39:01 deploy.go:391: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH]
|
||||
2026/09/25 10:40:28 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2
|
||||
2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for docmost → /mnt/sys_drive/felhom-data/backups/primary/docmost (images=3, secrets-referenced=2, data_keys=0, portable-carried=2/2, with
|
||||
2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2
|
||||
2026/09/25 10:46:25 handlers.go:1707: [WARN] [web] Restore requested (async): stack=nextcloud, snapshot=helyi from 172.18.0.6:39110
|
||||
2026/09/25 10:46:25 restore_unit.go:306: [WARN] [backup] No readable recovery unit for nextcloud at /mnt/sys_drive/felhom-data/backups/primary/nextcloud — falling back to volume-only restore
|
||||
2026/09/25 10:46:25 restore.go:49: [INFO] [backup] Starting app-data restore for nextcloud (drive=/mnt/sys_drive)
|
||||
2026/09/25 10:46:27 manager.go:1488: [ERROR] [stacks] stderr: time="2026-09-25T10:46:25Z" level=warning msg="The \"HDD_PATH\" variable is not set. Defaulting to a blank string."
|
||||
stderr: time="2026-09-25T10:46:25Z" level=warning msg="The \"HDD_PATH\" variable is not set. Defaulting to a blank string."
|
||||
2026/09/25 10:46:27 restore.go:77: [WARN] RESTORE could not restart nextcloud after restore: starting stack nextcloud: exit code 1
|
||||
stderr: time="2026-09-25T10:46:25Z" level=warning msg="The \"HDD_PATH\" variable is not set. Defaulting to a blank string."
|
||||
2026/09/25 10:48:01 restore.go:91: [WARN] [backup] Restore completed but app health check failed: stack nextcloud did not reach running state within 1m30s after restore
|
||||
2026/09/25 10:48:01 restore.go:97: [INFO] RESTORE completed: stack=nextcloud
|
||||
2026/09/25 10:48:01 handlers.go:1719: [INFO] [web] Restore completed (async): stack=nextcloud in 1m35.743249387s (volumes 0/0, dbs 0/0)
|
||||
@@ -0,0 +1,12 @@
|
||||
before: deployed False state degraded
|
||||
('200', {'ok': True, 'message': 'Stack nextcloud stop completed'})
|
||||
remove (keep data, keep backups) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'verified': True}, 'message': 'Stack nextcloud removed'}
|
||||
Documents Nextcloud Manual.pdf Nextcloud intro.mp4 Nextcloud.png Photos Readme.md Reasons to use Nextcloud.pdf Templates Templates credits.md after-backup.txt before-backup.txt compose
|
||||
db-dumps
|
||||
manifest.json
|
||||
volume-dumps
|
||||
total 12
|
||||
drwxr-xr-x 3 root root 4096 Sep 15 09:17 .
|
||||
drwxr-xr-x 20 root root 4096 Sep 24 20:52 ..
|
||||
drwxr-xr-x 4 root root 4096 Sep 15 09:17 paperless
|
||||
|
||||
@@ -0,0 +1,28 @@
|
||||
# E1 Q2 — BY HAND: the unit's own volume tars poured into fresh volumes + the kept drive folder, started with the unit's HDD_PATH. 2026-09-25T11:00:44Z
|
||||
volume db_data restored from the unit's tar
|
||||
volume html restored from the unit's tar
|
||||
volume redis_data restored from the unit's tar
|
||||
Container nextcloud-db Healthy
|
||||
Container nextcloud Starting
|
||||
Container nextcloud Started
|
||||
- installed: true
|
||||
- version: 34.0.4.1
|
||||
- versionstring: 34.0.4
|
||||
- edition:
|
||||
13:01:14 nextcloud: readback of the seeded user found=True
|
||||
13:01:14 E1-Q2-hand: read nextcloud=True
|
||||
13:01:15 PROPFIND files of drilld112bd -> 207; before-backup.txt visible=True; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
|
||||
13:01:15 PROPFIND files of drilld112bd -> 207; after-backup.txt visible=False; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
|
||||
## occ files:scan --all
|
||||
Starting scan for user 2 out of 2 (drilld112bd)
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
| Folders | Files | New | Updated | Removed | Errors | Elapsed time |
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
| 11 | 130 | 0 | 4 | 0 | 0 | 00:00:00 |
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
took 0.6 s
|
||||
13:01:26 PROPFIND files of drilld112bd -> 207; after-backup.txt visible=True; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'after-backup.txt', 'before-backup.txt']
|
||||
## teardown of the hand-made stack
|
||||
Network nextcloud_nextcloud-internal Removing
|
||||
Network nextcloud_nextcloud-internal Removed
|
||||
0
|
||||
@@ -0,0 +1,21 @@
|
||||
# baseline on 0.272.0: install nextcloud over the kept folder (R-657)
|
||||
13:01:53 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud']
|
||||
13:01:53 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH']
|
||||
13:01:53 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
13:02:13 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.4-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'}
|
||||
True
|
||||
[Fri Sep 25 13:01:59.795525 2026] [mpm_prefork:notice] [pid 1:tid 1] AH00163: Apache/2.4.68 (Debian) PHP/8.5.10 configured -- resuming normal operations
|
||||
[Fri Sep 25 13:01:59.795564 2026] [core:notice] [pid 1:tid 1] AH00094: Command line: 'apache2 -D FOREGROUND'
|
||||
127.0.0.1 - - [25/Sep/2026:13:02:04 +0200] "GET /status.php HTTP/1.1" 200 1064 "-" "curl/8.14.1"
|
||||
127.0.0.1 - - [25/Sep/2026:13:02:34 +0200] "GET /status.php HTTP/1.1" 200 1064 "-" "curl/8.14.1"
|
||||
127.0.0.1 - - [25/Sep/2026:13:03:05 +0200] "GET /status.php HTTP/1.1" 200 1064 "-" "curl/8.14.1"
|
||||
- installed: true
|
||||
- version: 34.0.4.1
|
||||
- versionstring: 34.0.4
|
||||
|
||||
# Q3: remove with backups ALSO deleted, data kept
|
||||
backup paths the remove dialog offers: ['/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud']
|
||||
remove -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)'], 'backup_paths_removed': ['/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (928M)'], 'verified': True}, 'message': 'Stack nextcloud removed'}
|
||||
126M /mnt/felhom-drives/scratch_hdd/appdata/nextcloud
|
||||
Documents Nextcloud Manual.pdf Nextcloud intro.mp4 Nextcloud.png Photos Readme.md Reasons to use Nextcloud.pdf Templates Templates credits.md after-backup.txt before-backup.txt
|
||||
removed-app row for nextcloud on /backups/restore: False
|
||||
@@ -0,0 +1,51 @@
|
||||
# E1 — kept data: the spike on 9202 (controller 0.272.0, live catalog), 2026-09-25 12:36–13:04 CEST
|
||||
|
||||
Written before any build. nextcloud installed through the product with `HDD_PATH=/mnt/felhom-drives/scratch_hdd`
|
||||
(the drive root, as households have it — demo-hp's apps read `HDD_PATH: /mnt/felhom-drives/hdd_1`), so its files
|
||||
live at `<drive>/appdata/nextcloud`. Seeded through `occ user:add`; `before-backup.txt` written through WebDAV as that
|
||||
user; a unit taken by the product's own nightly db-dump leg (window moved to 12:44 by `POST /backups/window`, then back
|
||||
to 02:30 — `E1-01`, `E1-02`, `E1-03`); `after-backup.txt` written through WebDAV AFTER the unit.
|
||||
|
||||
## Q1 — remove „keep my data" (+ backups kept) → the removed-app row → restore: does nextcloud come back whole?
|
||||
|
||||
**FALSE on 0.272.0 — it comes back BROKEN.** (`E1-04`, `E1-05`)
|
||||
|
||||
- The row is there (R-487: `Nextcloud` in the `/backups/restore` picker), but the picker's snapshot list for it is
|
||||
EMPTY (`snapshots()` offered 0) — so a household cannot even choose a copy. The form post with `snapshot_id=helyi`
|
||||
(what the R-487 row sends) was accepted.
|
||||
- The restore then logged `No readable recovery unit for nextcloud at /mnt/sys_drive/felhom-data/backups/primary/
|
||||
nextcloud — falling back to volume-only restore` — **the unit sat on the DATA drive**, `.../scratch_hdd/backups/
|
||||
primary/nextcloud`, fully readable. `compose up` ran WITHOUT the app's env: `The "HDD_PATH" variable is not set`,
|
||||
the app container mounted `/appdata/nextcloud` on the guest's ROOT disk, `nextcloud-db` crash-looped
|
||||
(`Database is uninitialized and password option is not specified`), `volumes 0/0, dbs 0/0`, no app.yaml.
|
||||
- **Root cause (read from source):** `backup.primaryUnitDirFor` and `backup.ListRestorePoints` decide "is this app
|
||||
removed?" with `stackProvider.GetStackComposePath(name)`, which in production answers **true for every catalog
|
||||
app** (it asks whether the STACK exists — every template is a stack). The R-487 test's fake answers it for
|
||||
deployed apps only, so `TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits` passes while the box fails —
|
||||
a seam that lies. Neither R-487 branch has ever run on a box for a catalog app on a data drive.
|
||||
**Filed R-690, fixed in this build** (`isStackDeployed`, the list's own predicate); `/appdata/paperless` on
|
||||
9202's root (2026-09-15) looks like the same accident, earlier.
|
||||
|
||||
## Q2 — a file written after the last backup: visible, or does nextcloud need `occ files:scan`?
|
||||
|
||||
**TRUE — it needs the scan.** (`E1-07`, BY HAND — the product route was broken, so the unit's OWN volume tars were
|
||||
poured into fresh volumes and the stack started with the unit's `HDD_PATH`; the database is what the unit holds.)
|
||||
After the start: the seeded account reads back; `before-backup.txt` visible; **`after-backup.txt` NOT visible**
|
||||
though it is on the drive. `occ files:scan --all` (0.6 s, 130 files, 4 updated) → `after-backup.txt` visible.
|
||||
So nextcloud's template gets `after_load:` = `php occ files:scan --all` as `www-data` in `nextcloud`.
|
||||
|
||||
## Q3 — remove with the backups ALSO deleted → files only
|
||||
|
||||
**TRUE.** (`E1-08`) `remove_backups` deleted the unit (928 MB); `appdata/nextcloud` (126 MB, both files) stayed; the
|
||||
removed-app row is gone. So "use my kept data" must be OFF here (no database copy) — the brief's `use-off` sentence.
|
||||
|
||||
## And R-657's loop — measured again as a baseline
|
||||
|
||||
A fresh install over the kept folder on 0.272.0 did **not** loop this time: `occ status` → `installed: true` after
|
||||
60 s. The old account (`drilld112bd`) and its files are invisible to the new install — its database knows nothing of
|
||||
them. So the harm is data that silently stops being reachable, and the loop is one way it can show. Either way a
|
||||
silent install over kept files is what E2 replaces with the choice.
|
||||
|
||||
Teardown: nextcloud removed through the product (keep data, backups deleted); the hand-made stack `down -v`;
|
||||
`appdata/nextcloud` (126 MB) kept on purpose — it is E5's kept-data fixture. docmost untouched.
|
||||
Controller on 9202 was 0.272.0 throughout — not interrupted.
|
||||
@@ -0,0 +1,6 @@
|
||||
R-690 red-proof, 2026-09-25
|
||||
(1) primaryUnitDirFor with GetStackComposePath as the removed? test:
|
||||
FAIL: the restore opens <sys>/felhom-data/backups/primary/nextcloud, want the kept unit <drive>/backups/primary/nextcloud
|
||||
(2) ListRestorePoints with GetStackComposePath:
|
||||
FAIL: the picker offers [] (found=true), want the one kept copy on HDD
|
||||
Restored: ok
|
||||
@@ -0,0 +1,8 @@
|
||||
### RED-PROOF R-687 (1): copy the steps with append(nil, …)
|
||||
--- FAIL: TestR687_EmptyLegReportsEmptySteps (0.00s)
|
||||
unattended_test.go:473: an empty leg must report steps as [], got {"trigger":"after-offsite","started_at":"0001-01-01T00:00:00Z","ended_at":"0001-01-01T00:00:00Z","deadline":"0001-01-01T00:00:00Z","enabled":false,"done":0,"undone":0,"held":0,"failed":0,"skipped":0,"steps":null}
|
||||
FAIL
|
||||
### RED-PROOF R-687 (2): drop the taken-step log line
|
||||
--- FAIL: TestLeg_FilesMayChangeNeedsAWholeCopy (0.01s)
|
||||
unattended_test.go:250: a taken files_may_change step must log which whole copy allowed it
|
||||
FAIL
|
||||
@@ -0,0 +1,4 @@
|
||||
### RED-PROOF R-688: the preview drops cloudflare_manual
|
||||
--- FAIL: TestR688_PreviewNamesCloudflareByHand (0.03s)
|
||||
customer_delete_test.go:607: preview missing "\"cloudflare_manual\":["
|
||||
body: {"_x":["the Cloudflare tunnel that serves acme.example (this customer's config carried its own tunnel token)","the DNS records of acme.example (the apex and *.acme.example)"],"claim_present":true,"customer_id":"acme","customer_name":"Teszt","dr_recipe_present":true,"has_config":true,"host_count":1,"hosts":[{"host_id":"acme-01","online":false,"status":"pending"}],"offsite_enabled":false,"offsite_identifier":"","offsite_type":"","one_time_secret":true,"online_host_present":false,"pbs_tenancy_configured":false,"pending_journal":null,"residue":{"app_log_tails":0,"app_telemetry":0,"appliance_registrations":0,"log_tail_requests":0,"notification_prefs":0,"reports":0,"selfbind_tokens":0},"residue_total":0,"superseded_blobs":1}
|
||||
@@ -0,0 +1,9 @@
|
||||
A-seed: seed docmost=True
|
||||
A-read0: read docmost=True
|
||||
A4-after-undo: read docmost=False
|
||||
A4-after-undo-retry: read docmost=True
|
||||
A4-after-undo-productstart: read docmost=True
|
||||
E1-seed: seed nextcloud=True
|
||||
E1-seed2: seed nextcloud=True
|
||||
E1-Q1-after-restore: read nextcloud=False
|
||||
E1-Q2-hand: read nextcloud=True
|
||||
@@ -0,0 +1,7 @@
|
||||
# read 2026-09-25T10:16:49Z — debug ring, lines naming the night's legs
|
||||
5000 /var/lib/docker/volumes/felhom-controller-data/_data/data/debug-ring.log
|
||||
{"timestamp":"2026-09-25T09:46:12Z","level":"DEBUG","message":"[scheduler] job status-refresh: execution starting","source":"scheduler.go:67"}
|
||||
{"timestamp":"2026-09-25T09:46:12Z","level":"DEBUG","message":"[scheduler] job ring-spill: execution starting","source":"scheduler.go:67"}
|
||||
# app.yaml last_auto_update
|
||||
# settings app_update + backup window
|
||||
{'app_update': None, 'backup_window': None, 'backup': None}
|
||||
@@ -0,0 +1,7 @@
|
||||
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.359+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/felhom-golden-0.236.0.tar.zst landed=2026-09-13T20:17:01Z reason="newest settled archive (landed 2026-09-13T20:17:01Z) has not been proven (last proven archive was a different one)"
|
||||
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.375+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/felhom-golden-0.236.0.tar.zst err="reconcile: restore-test extract archive config: proxmox: GET /nodes/demo-hp/vzdump/extractconfig?volume=local%3Abackup%2Ffelhom-golden-0.236.0.tar.zst -> HTTP 403: permission denied at /vms/ (missing privilege VM.Backup)"
|
||||
-rw-r--r-- 1 root root 654115664 Sep 13 22:16 felhom-golden-0.236.0.tar.zst
|
||||
3
|
||||
Sep 24 10:36:30 demo-hp felhom-agent[1195783]: time=2026-09-24T10:36:30.331+02:00 level=ERROR msg="backup: scheduled res
|
||||
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.375+02:00 level=ERROR msg="backup: scheduled rest
|
||||
Sep 25 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T10:57:53.372+02:00 level=ERROR msg="backup: scheduled rest
|
||||
@@ -0,0 +1,64 @@
|
||||
# read-only, demo-hp 9201, GET /api/stacks/docmost 200
|
||||
deployed = true
|
||||
state = "running"
|
||||
catalog_images = {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}
|
||||
installed_images = {'docmost': {'ref': 'docmost/docmost:0.96.0', 'digest': 'sha256:b56947fcfd08aab8fae12a377e1792784786adbf8b96e4281f14ef4fc072685a', 'at': '2026-09-22T09:07:17Z'}, 'docmost-postgres': {'ref': 'postgres:16-alpine', 'digest': 'sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea', 'at': '2026-09-22T09:07:17Z'}, 'docmost-redis': {'ref': 'redis:7-alpine', 'digest': 'sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499', 'at': '2026-09-22T09:07:17Z'}}
|
||||
pinned_images = {'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
deployed_at = 2026-08-31T12:24:41Z
|
||||
['name', 'meta', 'compose_path', 'state', 'deployed', 'protected', 'orphaned', 'containers', 'app_config', 'deploying', 'updating', 'health_probe', 'last_updated', 'restarting_since', 'template_images', 'catalog_images', 'catalog_digests', 'catalog_tested_at']
|
||||
# GET /api/stacks/docmost/backup-data 200
|
||||
{
|
||||
"stack": "docmost",
|
||||
"backup_paths": [
|
||||
{
|
||||
"path": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
|
||||
"size_bytes": 134773579,
|
||||
"size_human": "129M",
|
||||
"exists": true
|
||||
},
|
||||
{
|
||||
"path": "/mnt/felhom-drives/hdd_1/backups/secondary/docmost",
|
||||
"size_bytes": 132154188,
|
||||
"size_human": "127M",
|
||||
"exists": true
|
||||
}
|
||||
],
|
||||
"has_backups": true
|
||||
}
|
||||
# units on the box (ls -la --time-style=full-iso)
|
||||
/mnt/sys_drive/felhom-data/backups/primary/docmost
|
||||
/mnt/felhom-drives/hdd_1/backups/secondary/docmost
|
||||
## /mnt/sys_drive/felhom-data/backups/primary/docmost
|
||||
total 24
|
||||
drwxr-xr-x 5 root root 4096 2026-09-25T09:31:43 .
|
||||
drwxr-xr-x 10 root root 4096 2026-09-23T12:33:30 ..
|
||||
drwxr-xr-x 2 root root 4096 2026-09-25T09:31:43 compose
|
||||
drwxr-xr-x 2 root root 4096 2026-09-25T02:15:01 db-dumps
|
||||
-rw-r--r-- 1 root root 1376 2026-09-25T09:31:43 manifest.json
|
||||
drwxr-xr-x 2 root root 4096 2026-09-25T02:15:37 volume-dumps
|
||||
{
|
||||
"schema_version": 2,
|
||||
"app_name": "docmost",
|
||||
"display_name": "Docmost",
|
||||
"controller_version": "0.272.0",
|
||||
"created_at": "2026-09-25T09:31:43Z",
|
||||
"drive": "/mnt/sys_drive",
|
||||
"namespace_root": "/mnt/sys_drive/felhom-data",
|
||||
"image_pins": [
|
||||
"docmost/docmost:0.96.0",
|
||||
"postgres:16-alpine",
|
||||
"redis:7-alpine"
|
||||
],
|
||||
"secret_env_vars": [
|
||||
"APP_SECRET",
|
||||
"DB_PASSWORD"
|
||||
],
|
||||
"data_key_env_vars": null,
|
||||
"secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT,
|
||||
## /mnt/felhom-drives/hdd_1/backups/secondary/docmost
|
||||
total 16
|
||||
drwxr-xr-x 3 root root 4096 2026-08-23T01:30:00 .
|
||||
drwxr-xr-x 9 root root 4096 2026-09-14T01:30:00 ..
|
||||
-rw-r--r-- 1 root root 1 2026-09-25T01:30:13 .felhom-tier2-layout
|
||||
drwxr-xr-x 5 root root 4096 2026-09-24T20:59:16 recovery-unit
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
# read 2026-09-25T10:16:45Z — debug ring, lines naming the night's legs
|
||||
5000 /var/lib/docker/volumes/felhom-controller-data/_data/data/debug-ring.log
|
||||
{"timestamp":"2026-09-25T05:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T05:49:14Z","level":"INFO","message":"[stacks] Status refresh: 4 containers across 56 stacks","source":""}
|
||||
{"timestamp":"2026-09-25T05:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T05:54:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T05:54:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T05:59:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T05:59:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:04:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:04:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:09:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:09:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:14:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:14:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:19:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:19:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:24:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:24:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:29:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:29:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:34:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:34:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:39:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:39:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:44:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:44:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:49:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:54:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:54:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T06:59:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T06:59:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:04:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:04:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:09:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:09:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:14:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:14:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:19:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:19:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:24:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:24:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:29:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:29:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:34:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:34:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:39:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:39:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:44:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:44:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:49:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:54:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:54:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
{"timestamp":"2026-09-25T07:59:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
|
||||
{"timestamp":"2026-09-25T07:59:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
|
||||
# app.yaml last_auto_update
|
||||
/opt/docker/stacks/opengist/app.yaml
|
||||
last_auto_update:
|
||||
at: "2026-09-25T02:15:46Z"
|
||||
outcome: done
|
||||
from:
|
||||
# settings app_update + backup window
|
||||
{'app_update': None, 'backup_window': None, 'backup': None}
|
||||
@@ -0,0 +1,5 @@
|
||||
# read from a copy of the hub DB (+wal), 2026-09-25 ~12:30 CEST; the copy is deleted after
|
||||
demo-felhom first seen in report received 2026-09-25 02:29:15 update_leg= {"trigger": "after-offsite", "started_at": "2026-09-25T02:15:46.825918487Z", "ended_at": "2026-09-25T02:16:06.961273956Z", "deadline": "2026-09-25T07:30:00+02:00", "enabled": true, "done": 1, "undone": 0, "held": 0, "failed": 0, "skipped": 0, "steps": [{"app": "opengist", "outcome": "done", "from": {"opengist": "ghcr.io/thomiceli/opengist:1.13"}, "to": {"opengist": "ghcr.io/thomiceli/opengist:1.15"}, "seconds": 20.1}]}
|
||||
demo-felhom latest report 2026-09-25 10:16:42 controller 0.272.0
|
||||
demo-hp first seen in report received 2026-09-25 02:29:17 update_leg= {"trigger": "after-offsite", "started_at": "2026-09-25T02:18:41.05475777Z", "ended_at": "2026-09-25T02:18:41.126639901Z", "deadline": "2026-09-25T07:30:00+02:00", "enabled": true, "done": 0, "undone": 0, "held": 0, "failed": 0, "skipped": 0, "steps": null}
|
||||
demo-hp latest report 2026-09-25 10:16:43 controller 0.272.0
|
||||
@@ -0,0 +1,47 @@
|
||||
== felhom-pve
|
||||
Sep 24 21:34:05 demo-felhom felhom-agent[3559458]: time=2026-09-24T21:34:05.465+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-felhom-8363b5 guest="" source=one_shot_job disposition=destructive allowed=true reason=
|
||||
Sep 24 21:34:05 demo-felhom felhom-agent[3559458]: time=2026-09-24T21:34:05.465+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
|
||||
Sep 24 21:34:09 demo-felhom felhom-agent[3858230]: time=2026-09-24T21:34:09.787+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_w
|
||||
Sep 24 21:36:37 demo-felhom felhom-agent[3860787]: time=2026-09-24T21:36:37.184+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 24 21:36:38 demo-felhom felhom-agent[3860787]: time=2026-09-24T21:36:38.416+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_w
|
||||
Sep 24 22:51:41 demo-felhom felhom-agent[3860787]: time=2026-09-24T22:51:41.639+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-felhom-8363b5 guest="" source=one_shot_job disposition=destructive allowed=true reason=
|
||||
Sep 24 22:51:41 demo-felhom felhom-agent[3860787]: time=2026-09-24T22:51:41.639+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
|
||||
Sep 24 22:51:44 demo-felhom felhom-agent[3925782]: time=2026-09-24T22:51:44.757+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 24 22:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-24T22:51:45.908+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_w
|
||||
Sep 25 04:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-25T04:51:45.379+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-22T04:12:20Z) is already prov
|
||||
Sep 25 07:30:31 demo-felhom felhom-agent[3925782]: time=2026-09-25T07:30:31.104+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-backup archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_25-07_29_19.tar.zst size_bytes=2752
|
||||
Sep 25 07:30:31 demo-felhom felhom-agent[3925782]: time=2026-09-25T07:30:31.104+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-backup job=backup-9201-1790314159465759090 archive=felhom-backup:backup/vzdump-lxc
|
||||
Sep 25 10:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:51:45.064+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-backup archive=felhom-backup:backup/vzd
|
||||
Sep 25 10:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:51:45.909+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
|
||||
Sep 25 10:52:58 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:52:58.751+02:00 level=INFO msg="audit: gate decision" class=guest_destroy host=demo-felhom-8363b5 guest=990000 source=one_shot_job disposition=benign allowed=true reason=
|
||||
Sep 25 10:52:58 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:52:58.751+02:00 level=INFO msg="gate decision" class=guest_destroy guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign
|
||||
Sep 25 10:53:04 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:53:04.805+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_24-07_27_27.tar.zst duration_s=79.735774907
|
||||
-rw-r--r-- 1 root root 5840583923 2026-07-28_06:49:30 vzdump-lxc-9201-2026_07_28-06_44_22.tar.zst
|
||||
-rw-r--r-- 1 root root 16 2026-07-28_06:49:30 vzdump-lxc-9201-2026_07_28-06_44_22.tar.zst.notes
|
||||
-rw-r--r-- 1 root root 1579 2026-07-28_17:58:35 vzdump-lxc-9201-2026_07_28-17_53_26.log
|
||||
-rw-r--r-- 1 root root 5929975285 2026-07-28_17:58:34 vzdump-lxc-9201-2026_07_28-17_53_26.tar.zst
|
||||
-rw-r--r-- 1 root root 16 2026-07-28_17:58:34 vzdump-lxc-9201-2026_07_28-17_53_26.tar.zst.notes
|
||||
== demo-hp
|
||||
Sep 24 21:48:57 demo-hp felhom-agent[3717779]: time=2026-09-24T21:48:57.876+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-hp-bb76ea guest="" source=one_shot_job disposition=destructive allowed=true reason=signed k
|
||||
Sep 24 21:48:57 demo-hp felhom-agent[3717779]: time=2026-09-24T21:48:57.876+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
|
||||
Sep 24 21:49:02 demo-hp felhom-agent[728032]: time=2026-09-24T21:49:02.430+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
|
||||
Sep 24 21:56:35 demo-hp felhom-agent[752235]: time=2026-09-24T21:56:35.100+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 24 21:56:35 demo-hp felhom-agent[752235]: time=2026-09-24T21:56:35.952+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
|
||||
Sep 24 21:57:44 demo-hp felhom-agent[756657]: time=2026-09-24T21:57:44.968+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 24 21:57:45 demo-hp felhom-agent[756657]: time=2026-09-24T21:57:45.838+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
|
||||
Sep 24 22:06:23 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:23.065+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst size_bytes=8182759056 uncovered_volu
|
||||
Sep 24 22:06:23 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:23.065+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790279965190292701 archive=local:backup/vzdump-lxc-9201-2026_09_24-21_5
|
||||
Sep 24 22:07:45 demo-hp felhom-agent[756657]: time=2026-09-24T22:07:45.839+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=backup:felhom-pbs
|
||||
Sep 24 22:57:49 demo-hp felhom-agent[756657]: time=2026-09-24T22:57:49.433+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-hp-bb76ea guest="" source=one_shot_job disposition=destructive allowed=true reason=signed ke
|
||||
Sep 24 22:57:49 demo-hp felhom-agent[756657]: time=2026-09-24T22:57:49.433+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
|
||||
Sep 24 22:57:53 demo-hp felhom-agent[995644]: time=2026-09-24T22:57:53.035+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 24 22:57:53 demo-hp felhom-agent[995644]: time=2026-09-24T22:57:53.898+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
|
||||
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.359+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/felhom-golden-0.236.0.ta
|
||||
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.375+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/felhom-golden-0.236.0.tar.zst err="reconcile: restore-test extract archive config:
|
||||
Sep 25 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T10:57:53.357+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/felhom-golden-0.236.0.ta
|
||||
Sep 25 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T10:57:53.372+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/felhom-golden-0.236.0.tar.zst err="reconcile: restore-test extract archive config:
|
||||
-rw-r--r-- 1 root root 187 2026-09-24_08:22:55 vzdump-lxc-9201-2026_09_24-08_22_55.log
|
||||
-rw-r--r-- 1 root root 1494 2026-09-24_22:06:19 vzdump-lxc-9201-2026_09_24-21_59_25.log
|
||||
-rw-r--r-- 1 root root 8182759056 2026-09-24_22:06:15 vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst
|
||||
-rw-r--r-- 1 root root 16 2026-09-24_22:06:15 vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst.notes
|
||||
-rw-r--r-- 1 root root 1780 2026-09-24_22:45:37 vzdump-lxc-9201-2026_09_24-22_40_21.log
|
||||
@@ -0,0 +1,57 @@
|
||||
"""d1.py <label> [--kill-on-empty] — ONE press of docmost's Update on 9202 (drill catalog), with the
|
||||
facts before and after: PG_VERSION (the engine's own file), the seed through docmost's front door,
|
||||
the phases with times, the controller's own conversion lines, the undo copies left. Evidence: D/<label>.txt."""
|
||||
import json, sys, subprocess, time, threading
|
||||
import walk as w, fixtures
|
||||
label = sys.argv[1]
|
||||
KILL = "--kill-on-empty" in sys.argv
|
||||
OUT = open(f"{w.EV}/D/{label}.txt", "w", buffering=1)
|
||||
def say(*a):
|
||||
w.say(*a); OUT.write(" ".join(map(str, a)) + "\n")
|
||||
TOK = json.load(open(w.SC + "/seed-tokens.json"))
|
||||
def pgver():
|
||||
return w.guest('docker exec docmost-postgres sh -c \'cat "$PGDATA/PG_VERSION"\' 2>&1; docker inspect docmost-postgres --format "{{.Config.Image}}" 2>&1').split()
|
||||
def seed():
|
||||
return fixtures.FIXTURES["docmost"].verify(w, "c-docs", TOK["docmost"], lambda m: None)
|
||||
w.login()
|
||||
w.sync_rescan("docmost", "postgres:18-alpine")
|
||||
st = w.stack("docmost")
|
||||
say(f"# {label} — {time.strftime('%Y-%m-%dT%H:%M:%S%z')}; controller {w.guest('cat /etc/felhom-controller-image').strip()}")
|
||||
say(f"before: PG_VERSION/image={pgver()} pinned={(st.get('app_config') or {}).get('pinned_images')} catalog={st.get('catalog_images')}")
|
||||
say(f"before: seed reads back through the front door = {seed()}")
|
||||
since = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip()
|
||||
killer = None
|
||||
if KILL:
|
||||
def kill():
|
||||
# wait INSIDE the guest for the controller's own "emptied" line, then SIGKILL its main process
|
||||
out = w.guest(f"""timeout 600 sh -c 'docker logs -f --since {since} felhom-controller 2>&1 | grep -m1 "emptied"'
|
||||
pid=$(docker inspect felhom-controller --format '{{{{.State.Pid}}}}'); kill -9 $pid; echo "KILLED controller pid $pid at $(date -u +%T.%N)" """, timeout=700)
|
||||
say(" [kill] " + out.strip().replace("\n", " | "))
|
||||
killer = threading.Thread(target=kill); killer.start(); time.sleep(2)
|
||||
r = w.press_update("docmost", poll=0.5, cap_s=1500)
|
||||
if killer:
|
||||
killer.join()
|
||||
say(" [kill] waiting for the controller to come back and the journal to be resumed")
|
||||
for i in range(120):
|
||||
time.sleep(5)
|
||||
try:
|
||||
w.login(); s2 = w.stack("docmost")
|
||||
except SystemExit:
|
||||
continue
|
||||
if not s2.get("updating") and s2.get("update_phase") in ("undone", "failed", "done"):
|
||||
r["after_restart"] = {"phase": s2.get("update_phase"), "error": s2.get("update_error"), "hold": s2.get("hold_reason")}
|
||||
break
|
||||
for p in r.get("phases", []):
|
||||
OUT.write(f" phase +{p['t']}s {p['phase']} | {p['label']} | err={p['error']}\n")
|
||||
say(f"result: {json.dumps({k: v for k, v in r.items() if k != 'phases'}, ensure_ascii=False)[:900]}")
|
||||
for i in range(40):
|
||||
if w.app_curl("c-docs", "/")[1] == "200":
|
||||
break
|
||||
time.sleep(5)
|
||||
st = w.stack("docmost")
|
||||
say(f"after: PG_VERSION/image={pgver()} pinned={(st.get('app_config') or {}).get('pinned_images')} conversion_copy={(st.get('app_config') or {}).get('conversion_copy')}")
|
||||
say(f"after: seed reads back through the front door = {seed()}")
|
||||
say("controller lines (docmost / conversion):")
|
||||
OUT.write(w.guest(f"docker logs --since {since} felhom-controller 2>&1 | grep -E 'update docmost|CONVERT|pg_dumpall|emptied|loaded into|resum|UNDO' | cut -c1-400") + "\n")
|
||||
OUT.write("undo copies now: " + w.guest("docker volume ls -q --filter label=felhom.undo-copy-of=docmost | tr '\\n' ' '") + "\n")
|
||||
say("done")
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,19 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Part C step 1 on 9202: move the backup window forward (the product's own form, POST /backups/window) so the
|
||||
night chain runs now: db-dump at W, Tier 2 at W+60, off-site at W+105 (9202 has none), gate at W+2h."""
|
||||
import sys, time, json
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
W = sys.argv[1]
|
||||
w.login()
|
||||
before = w.guest("docker logs --since 20h felhom-controller 2>&1 | grep 'Daily job' | tail -3").strip()
|
||||
w.say("before:\n" + before)
|
||||
sess = open(f"{w.SC}/sess{__import__('os').getpid()}.txt").read().strip()
|
||||
csrf = open(f"{w.SC}/csrf{__import__('os').getpid()}.txt").read().strip()
|
||||
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", f"window_start={W}", f"{w.BASE}/backups/window"])
|
||||
w.say(f"POST /backups/window window_start={W} -> {r.stdout}")
|
||||
time.sleep(3)
|
||||
after = w.guest("docker logs --since 2m felhom-controller 2>&1 | grep -E 'Daily job|window' | tail -6").strip()
|
||||
w.say("after:\n" + after)
|
||||
open(sys.argv[2], "w").write("\n".join(w.LOG) + "\n")
|
||||
@@ -0,0 +1,15 @@
|
||||
"""ncfile.py put|ls <name> — a file through Nextcloud's OWN WebDAV as the seeded user (front door)."""
|
||||
import json, sys, subprocess, re
|
||||
import walk as w
|
||||
t = json.load(open(w.SC + "/seed-tokens.json"))["nextcloud"]
|
||||
mode, name = sys.argv[1], sys.argv[2]
|
||||
base = f"https://192.168.0.114/remote.php/dav/files/{t['uid']}/"
|
||||
a = ["curl", "-sk", "-u", f"{t['uid']}:{t['pw']}", "-H", "Host: c-nc.enkisfelhom.hu", "-w", "\n%{http_code}"]
|
||||
if mode == "put":
|
||||
r = subprocess.run(a + ["-X", "PUT", "--data-binary", f"E1 {name}", base + name], capture_output=True, text=True)
|
||||
w.say(f"PUT {name} -> {r.stdout.strip().splitlines()[-1]}")
|
||||
else:
|
||||
r = subprocess.run(a + ["-X", "PROPFIND", "-H", "Depth: 1", base], capture_output=True, text=True)
|
||||
code = r.stdout.strip().splitlines()[-1]
|
||||
names = sorted(set(re.findall(r"<d:href>[^<]*/([^/<]+)</d:href>", r.stdout)))
|
||||
w.say(f"PROPFIND files of {t['uid']} -> {code}; {name} visible={name in names}; listing={names}")
|
||||
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
|
||||
|
||||
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
|
||||
`controller.yaml.pre-night0926` (NOT the older `.pre-28`, which a restore must never pick up).
|
||||
"""
|
||||
import re, sys, io
|
||||
sys.path.insert(0, '.')
|
||||
import walk as w
|
||||
|
||||
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
|
||||
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
|
||||
|
||||
|
||||
def creds():
|
||||
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
|
||||
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
|
||||
if m:
|
||||
return m.group(1), m.group(2)
|
||||
raise SystemExit("no admin credential")
|
||||
|
||||
|
||||
def to_drill():
|
||||
u, t = creds()
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
test -f {VOL}/controller.yaml.pre-night0926 || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-night0926
|
||||
python3 - <<'PY'
|
||||
import re
|
||||
p = "{VOL}/controller.yaml"
|
||||
s = open(p).read()
|
||||
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
|
||||
if not re.search(r'^update:', s, re.M):
|
||||
s += "update:\\n health_timeout: 90s\\n"
|
||||
open(p, "w").write(s)
|
||||
PY
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -A2 '^update:' {VOL}/controller.yaml
|
||||
"""))
|
||||
|
||||
|
||||
def restore():
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
cp -p {VOL}/controller.yaml.pre-night0926 {VOL}/controller.yaml
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -c '^update:' {VOL}/controller.yaml || true
|
||||
"""))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
to_drill() if sys.argv[1] == "drill" else restore()
|
||||
@@ -0,0 +1,24 @@
|
||||
"""seed | read — the four drill apps' data through their OWN front doors (fixtures.py, R-156). Tokens are kept in the
|
||||
scratchpad (they can carry generated credentials); the evidence gets only True/False per app."""
|
||||
import json, sys, os
|
||||
import walk as w, fixtures
|
||||
SUBS = {"wishlist": "c-wish", "navidrome": "c-navi", "romm": "c-romm", "vikunja": "c-vik"}
|
||||
TOK = w.SC + "/seed-tokens.json"
|
||||
mode, label = sys.argv[1], sys.argv[2]
|
||||
w.login()
|
||||
res = {}
|
||||
if mode == "seed":
|
||||
toks = {}
|
||||
for a, sub in SUBS.items():
|
||||
toks[a] = fixtures.FIXTURES[a].seed(w, sub, w.say)
|
||||
res[a] = toks[a] is not None
|
||||
json.dump(toks, open(TOK, "w"), default=str); os.chmod(TOK, 0o600)
|
||||
else:
|
||||
toks = json.load(open(TOK))
|
||||
for a, sub in SUBS.items():
|
||||
if len(sys.argv) > 3 and a not in sys.argv[3].split(","):
|
||||
continue
|
||||
res[a] = bool(toks.get(a)) and fixtures.FIXTURES[a].verify(w, sub, toks[a], w.say)
|
||||
line = f"{label}: {mode} " + " ".join(f"{a}={v}" for a, v in res.items())
|
||||
w.say(line)
|
||||
open(w.EV + "/C/data-readback.txt", "a").write(line + "\n")
|
||||
@@ -0,0 +1,19 @@
|
||||
"""sr.py seed|read <label> app=sub[,app=sub] — seed or read back through each app's OWN front door (fixtures.py).
|
||||
Tokens stay in the scratchpad; the evidence gets True/False only."""
|
||||
import json, sys, os
|
||||
import walk as w, fixtures
|
||||
mode, label = sys.argv[1], sys.argv[2]
|
||||
SUBS = dict(x.split("=") for x in sys.argv[3].split(","))
|
||||
TOK = w.SC + "/seed-tokens.json"
|
||||
w.login()
|
||||
toks = json.load(open(TOK)) if os.path.exists(TOK) else {}
|
||||
res = {}
|
||||
for a, sub in SUBS.items():
|
||||
if mode == "seed":
|
||||
toks[a] = fixtures.FIXTURES[a].seed(w, sub, w.say); res[a] = toks[a] is not None
|
||||
else:
|
||||
res[a] = bool(toks.get(a)) and fixtures.FIXTURES[a].verify(w, sub, toks[a], w.say)
|
||||
if mode == "seed":
|
||||
json.dump(toks, open(TOK, "w"), default=str); os.chmod(TOK, 0o600)
|
||||
line = f"{label}: {mode} " + " ".join(f"{a}={v}" for a, v in res.items())
|
||||
w.say(line); open(w.EV + "/data-readback.txt", "a").write(line + "\n")
|
||||
@@ -0,0 +1,29 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Night 2026-09-25 Part G (adapted from 09-24), the MACHINE layer on 9202: tonight's apps removed THROUGH THE PRODUCT, leftovers read by name,
|
||||
the backup window put back to 02:30 through the product's own form. Apps present before tonight stay:
|
||||
privatebin, paperless-ngx, filebrowser (gokapi was removed in chaos round 5 — R-644's disposition)."""
|
||||
import json, sys, time
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
w.login()
|
||||
apps = ["wishlist", "navidrome", "romm", "vikunja", "n8n", "mealie"]
|
||||
q = lambda: w.guest("for p in " + " ".join(apps) + "; do echo \"$p containers=$(docker ps -aq --filter label=com.docker.compose.project=$p | wc -l) volumes=$(docker volume ls -q | grep -c ^${p}_) undo_copies=$(docker volume ls -q --filter label=felhom.undo-copy-of=$p | wc -l)\"; done; docker ps --format '{{.Names}}' | sort | tr '\\n' ' '")
|
||||
before = q(); w.say("before:\n" + before)
|
||||
res = {}
|
||||
for a in apps:
|
||||
st = w.stack(a)
|
||||
if not st.get("deployed"):
|
||||
res[a] = "not deployed"; continue
|
||||
res[a] = w.remove(a)
|
||||
after = q(); w.say("after the product's removes:\n" + after)
|
||||
dirs = w.guest("ls -d /mnt/felhom-drives/scratch_hdd/userdata/*/ 2>/dev/null; ls /mnt/sys_drive/felhom-data/backups/secondary/ 2>/dev/null; ls -d /mnt/felhom-drives/scratch_hdd/userdata/*/backups/secondary/* 2>/dev/null")
|
||||
w.say("drive dirs left (read by name):\n" + dirs)
|
||||
sess = open(f"{w.SC}/sess{__import__('os').getpid()}.txt").read().strip()
|
||||
csrf = open(f"{w.SC}/csrf{__import__('os').getpid()}.txt").read().strip()
|
||||
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", "window_start=02:30", f"{w.BASE}/backups/window"])
|
||||
time.sleep(2)
|
||||
win = w.guest("docker logs --since 1m felhom-controller 2>&1 | grep -E 'rescheduled|window set' | tail -4")
|
||||
w.say(f"backup window -> 02:30: {r.stdout}\n{win}")
|
||||
json.dump({"before": before, "remove": res, "after": after, "dirs": dirs, "window": win}, open("../G/G2-machine-removes.json", "w"), indent=2)
|
||||
open("../G/G2-machine-removes.log", "w").write("\n".join(w.LOG) + "\n")
|
||||
@@ -0,0 +1,9 @@
|
||||
import sys, time, walk as w
|
||||
head = sys.argv[1]; w.login()
|
||||
for k in range(40):
|
||||
w.sync_rescan()
|
||||
got = w.guest("cd /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache && git rev-parse --short HEAD").strip()
|
||||
if got.startswith(head[:7]):
|
||||
print(f"box clone at {got} after {k+1} sync(s)"); sys.exit(0)
|
||||
time.sleep(8)
|
||||
sys.exit("box never caught up to " + head)
|
||||
@@ -0,0 +1,507 @@
|
||||
#!/usr/bin/env python3
|
||||
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
|
||||
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
|
||||
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
|
||||
and reads GET /api/stacks/<n>. No controller code exists for it.
|
||||
|
||||
The walk, per `09` §6.4 and the update-night brief §4:
|
||||
1 deploy from the DRILL catalog at the LIVE pin
|
||||
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
|
||||
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
|
||||
4 „Mentés most"
|
||||
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
|
||||
6 press the guarded Update, record every phase with timestamps
|
||||
7 read the seed back through the front door
|
||||
8 the four version observables side by side
|
||||
9 write the verdict record in `09`'s JSON shape
|
||||
|
||||
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
|
||||
"""
|
||||
import argparse, json, os, re, subprocess, sys, time
|
||||
from datetime import datetime, timezone
|
||||
|
||||
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/4e6cd5c6-d1fa-409b-ab25-405f6b8bb892/scratchpad"
|
||||
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/night-2026-09-26"
|
||||
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
|
||||
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
|
||||
GUEST = os.environ.get("GUEST", "9202")
|
||||
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
|
||||
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
|
||||
HOSTHDR = f"Host: felhom.{DOMAIN}"
|
||||
HP = "demo-hp"
|
||||
|
||||
LOG = []
|
||||
|
||||
|
||||
def say(*a):
|
||||
line = " ".join(str(x) for x in a)
|
||||
ts = datetime.now().strftime("%H:%M:%S")
|
||||
print(f"{ts} {line}", flush=True)
|
||||
LOG.append(f"{ts} {line}")
|
||||
|
||||
|
||||
def sh(args, timeout=300, inp=None):
|
||||
try:
|
||||
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
|
||||
except (subprocess.TimeoutExpired, OSError) as e:
|
||||
return subprocess.CompletedProcess(args, 124, "", f"{e}")
|
||||
|
||||
|
||||
def guest(script, timeout=600):
|
||||
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
|
||||
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
|
||||
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
|
||||
import secrets as _s
|
||||
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
|
||||
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
|
||||
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
|
||||
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
|
||||
timeout=timeout, inp=script)
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def login():
|
||||
pw = open(f"{SC}/.ctlpw").read().strip()
|
||||
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
|
||||
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
|
||||
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
|
||||
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
|
||||
if not m:
|
||||
sys.exit("login failed: no session cookie")
|
||||
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
|
||||
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
|
||||
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
|
||||
if not c:
|
||||
sys.exit("login failed: no csrf token")
|
||||
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
|
||||
|
||||
|
||||
def ctl(method, path, data=None, raw=False, tries=2):
|
||||
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
|
||||
in memory, so any controller restart during the night invalidates it silently."""
|
||||
for attempt in range(tries):
|
||||
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
|
||||
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
|
||||
if method != "GET":
|
||||
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
|
||||
"-X", method]
|
||||
if data is not None:
|
||||
args += ["--data", json.dumps(data)]
|
||||
args.append(f"{BASE}{path}")
|
||||
r = sh(args)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
if code.strip() in ("302", "401") and attempt + 1 < tries:
|
||||
login()
|
||||
continue
|
||||
if raw:
|
||||
return code.strip(), body
|
||||
try:
|
||||
return code.strip(), json.loads(body)
|
||||
except Exception:
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
|
||||
|
||||
def page(path):
|
||||
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
|
||||
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
|
||||
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
|
||||
"-w", "\n%{http_code}"]
|
||||
if method:
|
||||
args += ["-X", method]
|
||||
if data is not None:
|
||||
args += ["--data-binary", "@-"]
|
||||
args += list(extra) + [f"{BASE}{path}"]
|
||||
r = sh(args, timeout=timeout + 30, inp=data)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
return r.returncode, code.strip(), body
|
||||
|
||||
|
||||
def stack(name):
|
||||
_, d = ctl("GET", f"/api/stacks/{name}")
|
||||
return (d.get("data") or {}) if isinstance(d, dict) else {}
|
||||
|
||||
|
||||
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
|
||||
"""Settling says the container runs; this says the APP answers. Not the same thing."""
|
||||
last = None
|
||||
for _ in range(tries):
|
||||
rc, code, _ = app_curl(sub, path, timeout=15)
|
||||
last = (rc, code)
|
||||
if rc == 0 and code in want:
|
||||
return True
|
||||
time.sleep(delay)
|
||||
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
|
||||
return False
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the walk
|
||||
|
||||
|
||||
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
|
||||
|
||||
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
|
||||
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
|
||||
# password back off the box — the household sees it once. So the value the harness itself
|
||||
# generated is kept here for the life of the run, and nowhere else.
|
||||
GENERATED = {}
|
||||
|
||||
|
||||
def deploy_values(name, sub):
|
||||
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
|
||||
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
|
||||
|
||||
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
|
||||
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
|
||||
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
|
||||
reason this was safe to discover by running it (live-probes rule).
|
||||
|
||||
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
|
||||
the scratch drive first — the same act the drive browser performs for a household.
|
||||
"""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
|
||||
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
|
||||
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
|
||||
made = []
|
||||
for f in fields:
|
||||
ev, ty = f.get("env_var"), f.get("type")
|
||||
if ev in values:
|
||||
continue
|
||||
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
|
||||
# when the caller sends none, deliberately ("the user needs to know their password"),
|
||||
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
|
||||
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
|
||||
if not f.get("required") and ty != "password":
|
||||
continue # the controller generates the optional secrets itself
|
||||
if ty == "path":
|
||||
p = f"{DRIVE}/{name}"
|
||||
values[ev] = p
|
||||
made.append(p)
|
||||
elif ty in ("secret", "password"):
|
||||
import secrets as _s
|
||||
values[ev] = "Drill-" + _s.token_hex(12)
|
||||
GENERATED.setdefault(name, {})[ev] = values[ev]
|
||||
elif f.get("default"):
|
||||
values[ev] = f["default"]
|
||||
else:
|
||||
values[ev] = f"drill-{name}"
|
||||
if made:
|
||||
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
|
||||
say(f" [1] made the drive paths this app requires: {made}")
|
||||
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
|
||||
if extra:
|
||||
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
|
||||
return values
|
||||
|
||||
|
||||
def deploy(name, sub, extra_values=None):
|
||||
st = stack(name)
|
||||
if st.get("deployed"):
|
||||
say(f" [1] {name} already deployed — reusing")
|
||||
return True
|
||||
values = deploy_values(name, sub)
|
||||
if extra_values:
|
||||
values.update(extra_values)
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
|
||||
say(f" [1] deploy -> {code} {str(d)[:120]}")
|
||||
if code != "202":
|
||||
return False
|
||||
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
|
||||
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
|
||||
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
|
||||
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
|
||||
seen = None
|
||||
for _ in range(90):
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
seen = st.get("state")
|
||||
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
|
||||
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
|
||||
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
|
||||
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
|
||||
# (`runComposeDeploy` writes it), so that is what to wait for.
|
||||
pins = (st.get("app_config") or {}).get("pinned_images")
|
||||
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
|
||||
say(f" [1] deployed, controller state={seen}, "
|
||||
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
|
||||
if seen != "running":
|
||||
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
|
||||
f"not treated as a failure; the fixture's front-door wait is the real gate")
|
||||
return True
|
||||
say(f" [1] never became deployed (last controller state={seen!r})")
|
||||
return False
|
||||
|
||||
|
||||
def backup_now(name):
|
||||
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
|
||||
|
||||
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
|
||||
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
|
||||
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
|
||||
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
|
||||
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
|
||||
phase list. A seed written "after the backup" is therefore written before the update's own backup
|
||||
— the undo's last-second copy is still the one that must bring it back."""
|
||||
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
|
||||
return None
|
||||
|
||||
def drill_bump(app, frm, to, service_hint=None):
|
||||
"""Serialised across concurrent walks: one git working tree, one lock."""
|
||||
import fcntl
|
||||
with open(f"{SC}/drill.lock", "w") as lk:
|
||||
fcntl.flock(lk, fcntl.LOCK_EX)
|
||||
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
|
||||
return _drill_bump(app, frm, to, service_hint)
|
||||
|
||||
|
||||
def _drill_bump(app, frm, to, service_hint=None):
|
||||
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
|
||||
|
||||
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
|
||||
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
|
||||
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
|
||||
every service that carries the app's own version and no others.
|
||||
"""
|
||||
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
|
||||
fy = f"{DRILL}/templates/{app}/.felhom.yml"
|
||||
s = open(comp).read()
|
||||
froms = [x.strip() for x in frm.split(",") if x.strip()]
|
||||
tos = [x.strip() for x in to.split(",") if x.strip()]
|
||||
if len(froms) != len(tos):
|
||||
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
|
||||
return None
|
||||
for f1, t1 in zip(froms, tos):
|
||||
if f"image: {f1}" not in s:
|
||||
say(f" [5] FROM ref not found in compose: {f1}")
|
||||
return None
|
||||
s = s.replace(f"image: {f1}", f"image: {t1}")
|
||||
open(comp, "w").write(s)
|
||||
f = open(fy).read()
|
||||
today = datetime.now().strftime("%Y-%m-%d")
|
||||
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
|
||||
open(fy, "w").write(f)
|
||||
sh(["git", "-C", DRILL, "add", "-A"])
|
||||
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
|
||||
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
|
||||
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
|
||||
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
|
||||
return h
|
||||
|
||||
|
||||
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
|
||||
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
|
||||
|
||||
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
|
||||
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
|
||||
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
|
||||
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
|
||||
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
|
||||
NUMBER R-607 asks for and has never had.
|
||||
"""
|
||||
t0 = time.time()
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(2)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
time.sleep(2)
|
||||
if not expect_app or not expect_ref:
|
||||
return None
|
||||
for i in range(tries):
|
||||
cat = stack(expect_app).get("catalog_images") or {}
|
||||
if expect_ref in cat.values():
|
||||
waited = round(time.time() - t0, 1)
|
||||
if i:
|
||||
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
|
||||
f"to {expect_ref} — R-607's window, measured")
|
||||
return waited
|
||||
time.sleep(delay)
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(1)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
|
||||
f"catalog_images = {stack(expect_app).get('catalog_images')}")
|
||||
return None
|
||||
|
||||
|
||||
def badges(name):
|
||||
out = {}
|
||||
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
|
||||
h = page(f"/apps/{name}{suffix}")
|
||||
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
|
||||
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
|
||||
return out
|
||||
|
||||
|
||||
def press_update(name, poll=1.0, cap_s=1800):
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/update")
|
||||
say(f" [6] Update -> {code} {str(d)[:220]}")
|
||||
if code not in ("202", "200"):
|
||||
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
|
||||
phases, seen, t0 = [], None, time.time()
|
||||
while time.time() - t0 < cap_s:
|
||||
st = stack(name)
|
||||
ph = st.get("update_phase")
|
||||
if ph != seen:
|
||||
seen = ph
|
||||
rec = {"t": round(time.time() - t0, 1), "phase": ph,
|
||||
"label": st.get("update_phase_label"), "updating": st.get("updating"),
|
||||
"error": st.get("update_error"), "hold": st.get("hold_reason")}
|
||||
phases.append(rec)
|
||||
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
|
||||
f"err={rec['error']} hold={rec['hold']}")
|
||||
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
|
||||
break
|
||||
time.sleep(poll)
|
||||
st = stack(name)
|
||||
return {"accepted": True, "http": code, "phases": phases,
|
||||
"duration_s": round(time.time() - t0, 1),
|
||||
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
|
||||
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
|
||||
|
||||
|
||||
def observables(name):
|
||||
st = stack(name)
|
||||
ac = st.get("app_config") or {}
|
||||
live = guest(f"""
|
||||
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
|
||||
echo '---inspect---'
|
||||
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
|
||||
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
|
||||
done
|
||||
""")
|
||||
a, _, b = live.partition("---inspect---")
|
||||
return {
|
||||
"pinned_images": ac.get("pinned_images"),
|
||||
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
|
||||
for k, v in (ac.get("installed_images") or {}).items()},
|
||||
"catalog_images": st.get("catalog_images"),
|
||||
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
|
||||
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
|
||||
}
|
||||
|
||||
|
||||
def app_logs(name, lines=400):
|
||||
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
|
||||
one string with escaped newlines — a scan over the envelope sees a single enormous line and
|
||||
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
|
||||
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
|
||||
if isinstance(d, dict):
|
||||
data = d.get("data")
|
||||
if isinstance(data, dict) and isinstance(data.get("logs"), str):
|
||||
return data["logs"]
|
||||
if isinstance(d.get("_raw"), str):
|
||||
return d["_raw"]
|
||||
return str(d)
|
||||
|
||||
|
||||
def write_verdict(rec, appdir):
|
||||
os.makedirs(appdir, exist_ok=True)
|
||||
p = os.path.join(appdir, "verdict.json")
|
||||
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
|
||||
say(f" [9] verdict {rec['verdict']} -> {p}")
|
||||
|
||||
|
||||
def remove(name):
|
||||
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
|
||||
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
|
||||
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
|
||||
say(f" [X] stop -> {c1} {str(d1)[:100]}")
|
||||
for _ in range(24):
|
||||
time.sleep(5)
|
||||
if stack(name).get("state") != "running":
|
||||
break
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": True, "remove_backups": True})
|
||||
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
|
||||
if code == "409":
|
||||
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
|
||||
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
|
||||
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
|
||||
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
|
||||
# harness takes it, then tidies its own directory by name at teardown.
|
||||
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
|
||||
" — removing the app and KEEPING the drive data instead")
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": False, "remove_backups": True})
|
||||
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
|
||||
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
|
||||
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
|
||||
return code
|
||||
|
||||
|
||||
def app_env(name, key):
|
||||
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
|
||||
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
|
||||
shows them the same value. Data still goes in through the app's own front door."""
|
||||
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
|
||||
if ":" in out:
|
||||
return out.split(":", 1)[1].strip().strip('"').strip("'")
|
||||
return ""
|
||||
|
||||
|
||||
def snapshots(name):
|
||||
"""The restorable copies the backups page offers for this app."""
|
||||
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
|
||||
data = d.get("data") if isinstance(d, dict) else None
|
||||
if isinstance(data, dict):
|
||||
for k in ("snapshots", "items", "restore_points"):
|
||||
if isinstance(data.get(k), list):
|
||||
return data[k]
|
||||
return data if isinstance(data, list) else []
|
||||
|
||||
|
||||
def restore(name, snapshot_id=None, wait_s=1200):
|
||||
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
|
||||
|
||||
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
|
||||
`snapshot_id` — because that is the button the sentence tells them to press.
|
||||
"""
|
||||
snaps = snapshots(name)
|
||||
if snapshot_id is None:
|
||||
if not snaps:
|
||||
say(f" [R] no restorable copy offered for {name}")
|
||||
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
|
||||
first = snaps[0]
|
||||
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
|
||||
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
|
||||
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
|
||||
"-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}",
|
||||
"--data-urlencode", f"stack_name={name}",
|
||||
"--data-urlencode", f"snapshot_id={snapshot_id}",
|
||||
f"{BASE}/backup/restore"], timeout=180)
|
||||
head = (r.stdout or "").split("\n")[0].strip()
|
||||
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
|
||||
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
|
||||
t0 = time.time()
|
||||
last = None
|
||||
while time.time() - t0 < wait_s:
|
||||
code, d = ctl("GET", "/api/backup/restore-status")
|
||||
dd = d.get("data") or {}
|
||||
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
|
||||
if cur != last:
|
||||
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
|
||||
last = cur
|
||||
if not dd.get("running", False) and time.time() - t0 > 5:
|
||||
break
|
||||
time.sleep(2)
|
||||
st = stack(name)
|
||||
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
|
||||
f"phase={st.get('update_phase')}")
|
||||
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
|
||||
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
|
||||
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
|
||||
"observables_after": observables(name)}
|
||||
@@ -675,7 +675,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. **-- RULED 2026-09-23 (`09` §3 decision 16):** PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it; the engine-major gate stays until then. | **READY TO BUILD — owner: CC; `09` §6.4; the gate stays until all eleven are proven** |
|
||||
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. **-- NARROWED 2026-09-25 (evening), catalog `6a4a5f0`, `09` §3 decision 35:** the PostgreSQL half now passes ONE app at a time — only a template whose ladder entry for the step is proven on BOTH venues and carries `engine_conversion` (the box converts it, controller v0.273.0), as the only image move in its commit. Every other PostgreSQL app stays refused; the postgis family is judged now (it was not). CLAUDE.md rule text updated the same commit. Decoys + red-proof: `audits/night-2026-09-26/C/`. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** |
|
||||
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). **PARTLY SHIPPED in v0.242.0 (`d698ce3`), measured live on 9202 the same night:** the difference is computed and a fresh compose-created volume IS reported (`["opengist_opengist_data"]`), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (`docker volume create <name>`, `restore.go:154`) carries no labels — compose still removes it and the response says `[]` (`audits/v0242-2026-09-14/19-R489-cause.txt`). **Remaining fix:** list by the `<project>_` name prefix as well (union), or label the recreated volume as compose would. | **READY — rank P3-LOW; owner: CC (residual)** |
|
||||
| **R-492** | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** |
|
||||
@@ -810,6 +810,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` | **READY — P3; owner: CC (hub) / operator (Peti's Cloudflare leftovers, if any)** |
|
||||
| **R-689** | **[P3-LOW] demo-hp's scheduled restore test picks the golden template in `local:backup/` as a backup and fails every 6 h.** Read 2026-09-25 (agent 0.134.0): `restore-test tier is DUE … archive=local:backup/felhom-golden-0.236.0.tar.zst … newest settled archive (landed 2026-09-13T20:17:01Z)` then `scheduled restore-test FAILED … extractconfig … HTTP 403: permission denied at /vms/ (missing privilege VM.Backup)` — 2026-09-24 10:36, 2026-09-25 04:57 and 10:57. The golden is not a backup of any guest; the real archive (`vzdump-lxc-9201-2026_09_24-21_59_25`) was younger than the 24 h settle. So the restore test proves nothing on demo-hp and logs an ERROR each cycle. **Fix direction:** the restore test selects only `vzdump-*` archives of a known guest (or the golden moves out of `local:backup/`). Agent-side; not built (the agent is untouched this session). `audits/night-2026-09-26/part0/P0-demo-hp-restoretest-golden.txt` | **READY — P3; owner: CC (agent)** |
|
||||
| **R-690** | **[P1-HIGH] The removed-app restore (R-487) never finds a unit kept on a DATA drive: nextcloud came back with no env, no database and its files mounted on the guest's root disk.** MEASURED 2026-09-25 on 9202 (controller 0.272.0): nextcloud removed with „keep my data" and backups kept; the `/backups/restore` picker listed it but offered 0 copies; the restore logged `No readable recovery unit for nextcloud at /mnt/sys_drive/... — falling back to volume-only restore` while the unit sat readable on the data drive; `compose up` without env (`HDD_PATH` blank → bind `/appdata/nextcloud` on the root disk), `nextcloud-db` crash-looping, `volumes 0/0, dbs 0/0`. **Cause:** `backup.primaryUnitDirFor` and `ListRestorePoints` ask `GetStackComposePath` whether the app is removed; production answers true for EVERY catalog app (the stack exists), while the R-487 test's fake answers it for deployed apps only — the test passes, the box fails. **Fix:** ask `isStackDeployed` (the list's own predicate), pinned by a test on a provider that answers like production. `audits/night-2026-09-26/E/E1-README.md` Q1 | **FIX BUILT — controller (Part E build, 2026-09-25); live proof pending the release** — owner CC |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user