Files
felhom.eu/documentation/audits/SPIKE-storagebox-restic-direct-2026-07-09.md
T
admin 9b250b97c9 docs(audit): restic→Storage Box direct-from-residential spike (FIT/IPv4)
Validated the controller offbox restic-SFTP tier direct from the residential
demo box (IPv4-only egress) end-to-end through the real endpoints: reach on
:23 (PASS), chown-immune restic init (works where PBS-direct failed),
config→run→restore green, sub-account isolation confirmed. Residential upload
~4.3 MB/s. Concentrator NOT needed. KEY BLOCKER F1: offbox reads recovery unit
from the app data-drive namespace root, not the actual backup-drive location →
silently backs up 0 apps (ok status) on multi-drive setups. Spike only,
synthetic data, torn down.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-09 11:05:36 +02:00

211 lines
14 KiB
Markdown

# SPIKE — restic → Hetzner Storage Box, DIRECT from a residential (IPv4) customer box
**Date:** 2026-07-09 **Executor:** Claude Code (live) **Class:** spike (empirical validation)
**Egress under test:** the real `felhom-controller` container in demo LXC guest 9201 on `felhom-pve`
(residential N100 — the true customer-egress profile). **No controller code changed; no version bump.**
---
## 1. VERDICT
**Direct restic-over-SFTP from a residential IPv4 customer box → FIT.** The reachability that matters
for the pilot is proven: **the box is reachable on :23 over IPv4** from the residential egress, and the
controller's existing offbox tier drives a full **config → run → restore** loop against the live box
through its **real endpoints**. Sub-account-per-customer **isolation holds**. A **concentrator (ep0/WG)
is NOT required** for reachability.
**One blocking caveat for productionizing — a controller bug, not a box/restic problem (see F1):** the
offbox runner looks for each app's recovery unit under the app's **data-drive** namespace root, but on
any system whose primary backups live on a **separate backup drive** (this demo → the USB drive) it
finds nothing, **skips every app, and still reports `LastStatus=ok` with 0 snapshots**. Direct offbox is
mechanically sound; this path-resolution mismatch must be fixed (or the storage layout aligned) before it
backs up real customer data.
---
## 2. Verified inputs (live)
| Fact | Value |
|------|-------|
| Box | BX11 `PBS-storage-1` (#611421), `u629193`, **port 23**, FSN1, 1 TB, €4.06/mo, 0 B used at start |
| Main host | `u629193.your-storagebox.de`**now resolves A `167.235.97.81`** (AAAA `2a01:4f8:2b01:d2d::2` exists but see §3) |
| External reachability | **ON** (was OFF in the prior ep0 spike) |
| restic in container | **0.14.0** (go1.19.8) — uses `sftp.command`, already handled by offbox.go |
| ssh in container | OpenSSH 9.2p1 |
| Controller | v0.103.0, healthy, LXC 9201, container `felhom-controller`, listens `:8080` (bridge, in-container), dashboard open (no password) |
| ep0 | still holds a box-authorized key + a live SSHFS mount of `/home` (reused for teardown only) |
## 3. Reachability (Phase 1 — the gate) — **PASS over IPv4**
From **inside the controller container** (the residential egress):
- **DNS:** `getent ahostsv4``167.235.97.81`. `getent ahostsv6`**nothing**. The residential egress
is **IPv4-only** — no AAAA is returned to it, and an explicit `ssh -6` probe returns **"Network is
unreachable."** This is the **load-bearing result for the IPv4-only pilot (Peti).**
- **`ssh-keyscan -p 23`** returned host keys → reachable.
- **Explicit IPv4 probe** (`ssh -4 -p 23 … true`) → `Permission denied (publickey,password)` =
**reachable** (auth failure, not timeout/refused). Round-trip felt immediate.
- **Conclusion:** direct residential-IPv4 → box:23 works. No concentrator needed for reachability.
## 4. Host-key pinning (Phase 2 — no blind TOFU) — **PASS**
The three host keys seen from the residential container via `ssh-keyscan` are **byte-identical** to the
keys ep0 pinned over a **completely different network path** (intra-Hetzner) in the prior spike — two
independent paths agreeing rules out MITM. The verified line to pin in the controller config (ed25519,
fingerprint `SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM`):
```
[u629193.your-storagebox.de]:23 ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIICf9svRenC/PLKIL9nk6K/pxQgoiFC41wTNvoIncOxs
```
Note: the **sub-account hostname presents the SAME host key** (same physical box) — one pinned line
covers both main and sub endpoints.
## 5. Raw restic round-trip (Phase 3 — mechanism proof) — **PASS**
Reproduced offbox.go's exact `sftp.command` from inside the container, RepoPath `/home/felhom-spike`
(box `/` is read-only; `/home` writable), scratch key authorized on the box (see §Auth):
| Step | Result |
|------|--------|
| `restic init` | **OK — the chown-immunity headline.** restic writes plain files and never `chown`s, so it succeeds where **PBS-direct hit EACCES** (prior spike). This is *why* restic is the right tool for this box. |
| backup #1 (500 MB urandom) | **1:56 (116 s) → ~4.3 MB/s**; snapshot saved |
| restore latest + sha256 | **INTEGRITY-OK** (exact match) |
| backup #2 (identical) | **3 s, 704 B added** → dedup works over SFTP (factor 2.0) |
| `forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune` | completes, **32 s** |
### Auth path (Phase 3a)
Scratch ed25519 key generated in-container. Because ep0 already held a box-authorized key, the pubkey
was **appended to the box `authorized_keys` over SFTP using ep0's existing trust** — password-free, no
operator step (operator-approved alternative to `ssh-copy-id … -s`). Verified passwordless SFTP from the
residential container. (Observation: the box's main `authorized_keys` already held **11 keys** before
this — worth an operator cleanup.)
## 6. Integration proof through the REAL endpoints (Phase 4 — the proof that counts) — **PASS (with F1)**
Driven via curl to the in-container `:8080` (form POSTs via the public URL can silently no-op — a known
demo gotcha; localhost is authoritative). RepoPath `/home/felhom-offbox-spike`.
1. **`POST /backup/offbox/config`** (Host, Port 23, User `u629193`, RepoPath, `ssh_key` via
`--data-urlencode @file`, `known_hosts`, `enabled=on`) → 302 success flash. The controller wrote its
own secret set: `ssh_key` (0600), `known_hosts` (0644), **auto-generated `repo_password` (0600)**
fully controller-managed; the operator never supplies the repo password.
2. **`POST /backup/offbox/toggle`** (audiobookshelf) → included.
3. **`POST /backup/offbox/run`** → **F1 fired:** `LastStatus=ok`, **0 snapshots**, log:
`audiobookshelf: no backup data yet (/mnt/sys_drive/felhom-data/backups/primary/audiobookshelf) —
skipping`. The primary backups on this demo actually live on the **USB drive**
(`/mnt/felhom-drives/felhom-usb/backups/primary/…`); offbox looked under the app's **data-drive**
namespace root instead → found nothing.
4. After **staging the real (synthetic) recovery unit** at the path offbox expects, re-run →
**`LastStatus=ok`, 1 snapshot, 8 s.** Repo created correctly on the box (config/data/index/keys/
snapshots — a proper restic repo).
5. **`POST /backup/offbox/restore`** (audiobookshelf) → restored non-destructively to
`data/offbox-restore/audiobookshelf/…`; the full recovery unit (`compose/*.yml` + `manifest.json`)
came back and the **manifest sha256 matched the source exactly**.
**The config→run→restore loop is green through the customer/operator's real endpoints against the live
box over residential IPv4.**
### F1 (the important finding) — offbox recovery-unit source path vs actual backup location
`runOffboxInternal` derives the source as `RecoveryUnitPath(AppNamespaceRoot(stack), stack)` =
`<app-data-drive namespace root>/backups/primary/<stack>` (`internal/backup/offbox.go:268-276`,
`backup.go:147`). But primary backups are written to the configured **backup-storage drive**, which can
differ from the app's data drive. When they differ (this demo: data on `sys_drive`, backups on the
`felhom-usb` drive) offbox finds no recovery unit and skips. **Impact:** on any multi-drive customer
(likely Peti, if his backups target a dedicated drive) direct offbox would silently back up nothing.
**Fix direction:** resolve the recovery-unit source from the *actual primary-backup storage location*
(the backup target), not the app's data-drive namespace root — or guarantee the two are aligned.
*Confirm intended source with the controller team; this is the one real blocker for productionizing.*
### F2 — silent "ok / 0 snapshots"
A run that skips **every** app for missing recovery units still reports `LastStatus=ok`. There is no
warning surfaced to the operator/customer that nothing was backed up. Offbox should distinguish
"backed up N apps" from "0 apps had data" (warn/degraded status), else F1-class problems stay invisible.
## 7. Sub-account isolation (Phase 5) — **PASS**
Operator created sub-account **`u629193-sub1`** (own hostname `u629193-sub1.your-storagebox.de`, base
dir `/felhom-peti-spike`, SSH on). The **console SSH-key field did not take** (F3); the key was installed
with **`ssh-copy-id -p 23 -s -f -i <pub>`** using the sub-account password (both `-s` **and** `-f` are
required). Then, from the residential container:
- **Jail:** `ls /`*Permission denied*; the sub's `/home` shows **only its own `.ssh`** (its base dir
`/felhom-peti-spike` is presented as `/home`); **both main repos are "not found"** from the sub.
- **Round-trip:** init + backup + 1 snapshot with the sub's **distinct key + distinct repo password** — OK.
- **Isolation assert (b):** the **main** repo password on the **sub** repo → `Fatal: wrong password or no
key found` — repos are independently encrypted.
- **Isolation assert (a):** the sub endpoint pointed at the main repo path → *repository not found* —
the jail blocks cross-customer reads.
**Isolation model for fork 2:** sub-account-per-customer is viable. RepoPath convention **`/home/<name>`**
for both main and sub (the sub is chrooted, so `/home/<name>` is physically isolated per customer). User
pattern `u629193` (main) / `u629193-subN` (sub, with its own `…-subN.your-storagebox.de` hostname).
Provisioning is **console/Robot-driven**; the console SSH-key field is currently unreliable (F3) — use
`ssh-copy-id … -s -f` until fixed. `ValidateOffboxTarget` already accepts sub targets (hyphen in
user/host and the `/home/<name>` repo path all pass).
## 8. Throughput (Phase 6)
Residential upload is the sole bottleneck (intra-Hetzner box speed ~327 MB/s is irrelevant here). Clean
measurement: **500 MB in 116 s ≈ 4.3 MB/s (~34 Mbps upstream)**. Extrapolated first-seed:
**~20 GB ≈ 78 min, ~50 GB ≈ 195 min.** Restore is download-bound (faster on typical residential links).
**Implication:** the daily schedule is fine for steady-state deltas (dedup makes them tiny — 704 B/3 s
above), but the **first seed of a real app can take hours** — give the customer a heads-up on onboarding.
(A dedicated 5 GB timing was skipped as it only reconfirms the same upstream bottleneck.)
## 9. Provisioning values (bake into a future task)
| Field | Value |
|------|-------|
| Host | `u629193.your-storagebox.de` (main) · `u629193-subN.your-storagebox.de` (**sub has its own hostname**) |
| Port | **23** |
| User | `u629193` (main) · `u629193-subN` (sub) |
| RepoPath | **`/home/<name>`** (sub is chrooted → physically isolated) |
| known_hosts line | `[HOST]:23 ssh-ed25519 AAAA…Oxs` (same key for main + sub hostnames) |
| Key install | **`ssh-copy-id -p 23 -s -f -i <pub> USER@HOST`** (both `-s` and `-f` required; console SSH-key field unreliable per F3) |
| Retention | `--keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune` (offbox default) |
## 10. Secrets inventory for custody design (fork 4 — NOT decided here)
Per-customer, all controller-managed under `data/offbox/` (0600 key + repo_password, 0644 known_hosts):
- **restic repo password** — auto-generated by the controller on first config; **loss = unrecoverable
offsite data.**
- **SSH private key** — 0600.
Both ride DR via the PBS whole-CT snapshot. **Recovery-ordering coupling:** offsite-restic recovery
depends on first recovering these secrets from the PBS tier — the custody/escrow spec must sequence that.
## 11. Decision this unblocks
- **Topology (direct vs concentrator):** **DIRECT — recommended.** Residential IPv4 reaches the box
directly; no ep0/WG concentrator needed. The concentrator spike is **not** required.
- **Isolation (fork 2):** **sub-account-per-customer — recommended.** Jailed + independently encrypted.
Provisioning is console/Robot-driven (fix or work around the console SSH-key field, F3).
- **Relationship to the PBS tier:** restic offbox = the **app-data** offsite tier (backs up
`backups/primary/<stack>`), distinct from the PBS whole-CT tier. ep0 `felhom-offsite` capacity was
**not** re-observed in this spike (out of scope).
## 12. NOT done / next
- **F1 controller fix** — the recovery-unit source-path mismatch is the real blocker for direct offbox on
multi-drive customers; must be resolved (or storage layout aligned) before real data. **Highest priority.**
- **F2** — surface a warning when a run backs up 0 apps.
- **F3** — Hetzner console sub-account SSH-key field didn't apply the key; use `ssh-copy-id -s -f`.
- Custody/escrow design (fork 4) using the §10 inventory.
- Production provisioning task: small controller/UI polish + per-customer sub-account ceremony.
- Concentrator spike: **not needed** (direct works).
## 13. Teardown (left clean)
- **Box:** spike repos removed (`/home/felhom-spike`, `/home/felhom-offbox-spike`, sub `/home/peti-repo`);
main `authorized_keys` restored 12→**11** (spike key removed, ep0 key preserved); box `/home` back to
`.ssh` + the sub base dir.
- **Controller:** audiobookshelf un-toggled; staged recovery unit + `/mnt/sys_drive/felhom-data` removed;
`offbox-restore/` removed; `data/offbox/` secrets removed; **offbox target cleared from settings.json**;
controller restarted **healthy** (settings load clean, 0 offbox lines).
*(Mid-teardown a bad write truncated settings.json to 0 B; restored from the controller's own
`settings.json.bak` — no data lost. The daily `offbox-backup` job is still scheduled but no-ops with no config.)*
- **Container:** scratch keys/passwords (`/tmp/felhomspike`) removed.
- **OPERATOR TODO:** **delete (or rotate) sub-account `u629193-sub1`** — its password was shown on screen
during creation, so treat it as compromised scratch. Its base dir `/felhom-peti-spike/.ssh` (spike key)
goes away when the sub-account is deleted; keep it only if promoting to Peti's real sub-account (then
rotate the key). No secrets are recorded in this doc (all stored 0600 out-of-band).