Carolopedia

A public encyclopedia of Carolverse.

๐Ÿ“– Carolopedia โ€บ Guides โ€บ Infrastructure & Backups architectureGuide page

Infrastructure & Backups architecture

The operating runbook is not published.

๐Ÿ“–Summary

The Infrastructure & Backups service is built following the [agent-centric modular architecture](/dev/carolopedia/wiki/architecture) of Carolverse. It leverages agile principles to build or modify software using distinct agent identities, each carrying out a specific activity. Here, those activities protect every other service's data โ€” scheduled backups, retention and restores โ€” while keeping the runtime backbone that exposes each app alive and healthy.

๐Ÿ“–Functional considerations

This service is the safety net under every other service, so its architecture is shaped by what it must guarantee end-to-end:

- **Nothing is lost.** Every service's critical state is snapshotted on a schedule and shipped off the machine, so a single-host failure never loses data.
- **The right history is kept.** Retention keeps enough versions to recover from and prunes the rest, so backups protect without growing unbounded.
- **The machine stays healthy.** Disk, temp files, stale working files, logs and databases are tidied continuously so nothing silently fills up or rots.
- **Services stay reachable.** The systemd units, the web/proxy layer and the secure tunnel that expose each app are kept up; a down app is detected and relaunched rather than left dark.
- **Cost is attributable.** Protection is charged per protected data, so the service is accountable to a cost center.

๐Ÿ“–Solution architecture

The service is a set of **blocks** โ€” the distinct steps shown in the Blocks section of the service page โ€” each owned by an agent and carried out by that agent's droids. It is a direct instance of Carolverse's [agent-centric modular architecture](/dev/carolopedia/wiki/architecture).

- **Snapshot, then ship.** One path takes the snapshot of every service's data; a second pushes code commits offsite, with a dry-run twin that reviews push state before the real push.
- **Continuous housekeeping.** A separate, always-running cleanup path reclaims disk age-based and guards a disk-usage threshold, so backup and runtime never starve for space.
- **Self-healing runtime.** The app-steward path watches the service's registered apps and relaunches any that go down, rather than waiting for a human to notice.
- **Schedule-driven, not request-driven.** The work is recurring and time-triggered; there is no public request surface to protect data on demand.

๐Ÿ“–Technologies

- **Python 3** tooling for the backup, cleanup and app-steward droids, run on the Carol host behind **nginx**.
- **SQLite (WAL)** is what is being protected โ€” the snapshot captures the registry, design store, plan-generator, initiatives and constitution databases, plus laptop-critical assets.
- **Git / GitHub** is the offsite for code: unpushed commits are pushed so the remote is a durable copy.
- **systemd** and **cron** schedule the recurring droids (daily snapshot, daily janitor); **systemd units** plus a **secure tunnel** are the runtime layer that keeps each app reachable.
- The **registry** is the source of truth for which apps exist and must stay alive, and for cost-center attribution.

๐Ÿ“–Design principles

- **Single source of truth.** The set of apps to keep alive and the data to protect come from the live registry โ€” the shared principle described on the [Carolverse Architecture](/dev/carolopedia/wiki/architecture) page.
- **Offsite by default.** A backup that only lives on the same machine is not a backup; snapshots and commits are shipped off-host.
- **Self-heal over block-and-wait.** A down app is relaunched and stale disk is reclaimed automatically rather than escalated.
- **Agent-centric modular architecture.** Every block has an accountable agent and a doing droid.
- **Keep the right history, prune the rest.** Retention is deliberate, not unbounded growth.
- **Observability first.** Scheduled droids emit run-audit so a failed backup or cleanup is visible, not silent.

๐Ÿ“–Success criteria

- Every protected service's critical state has a **recent off-machine snapshot** that can be restored.
- **Unpushed commits do not accumulate** โ€” code is durably mirrored to GitHub.
- **Disk never silently fills** โ€” stale temp and caches are reclaimed and the disk-usage threshold holds.
- **Registered apps stay up** โ€” a down app is detected and relaunched without operator action.
- **Each recurring droid's runs are auditable**, so a missed or failed run surfaces on a monitor.

๐Ÿ“–Policies

- **Backups are scheduled, never ad-hoc** โ€” the snapshot and shipping run on their schedule, owned by the accountable agent.
- **Ship offsite** โ€” snapshots and commits must leave the host to count as protected.
- **Retention is enforced** โ€” keep the right window of history and prune the rest.
- **Every action is tagged to a droid** under the owning agent; recurring work emits run-audit so it is observable.
- **Protection is charged per protected data** against this service's cost center.

๐Ÿ“–What it delivers today

- **Daily off-machine snapshots** of Carol + BB database state (registry, designs, plan-generator, initiatives, constitution) plus laptop-critical assets โ€” Hagrid via the *Backup Custodian* droid (Backups & Shipping block).
- **Offsite code shipping** โ€” unpushed commits pushed to GitHub, with a dry-run review of push state first โ€” Hagrid via the *Shipper* and *Shipper Twin* droids.
- **Disk reclamation and guarding** โ€” pruning stale `/tmp` files and regenerable caches age-based, then guarding the disk-usage threshold โ€” Hagrid via the *Daily Janitor* droid (Cleanup & Housekeeping block).
- **App liveness** โ€” keeping the service's registered apps alive by detecting down apps and relaunching them via `shared machinery` โ€” Hagrid via the *App Steward* droid (Infrastructure & Runtime block).

๐Ÿ“–What it will deliver

- **Restore drills** โ€” periodic test restores from a snapshot to prove backups are recoverable, not just present.
- **Retention reporting** โ€” a clear view of what history exists per protected service and how it is pruned.
- **Per-service protection cost** surfaced against the cost center, so charges are visible alongside the data protected.

Source: Services Catalogue ยท Public information reflected here.