IaC + docs + failure ledger for the ai-services fleet
Find a file
VM460 dogfood fa4e3d10a4 ai-services-boilerplate: IaC, docs and failure ledger
Terraform/cloud-init/Ansible for VM 460 and the Windows pair, the ai-control
dashboard, the a2a-gateway, and docs/ including the compound-engineering
failure ledger, the Rust-port PRD and the agent mesh PRD.

Secrets are excluded by .gitignore (*.tfvars, .env, *.token); the Proxmox
token lives outside the repo at ~/.config/ai-services/pve.tfvars.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 17:25:10 +00:00
a2a-gateway ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
ansible ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
cloud-init ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
compose ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
control ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
docs ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
scripts ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
terraform ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
terraform-windows ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
.gitignore ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00
README.md ai-services-boilerplate: IaC, docs and failure ledger 2026-09-17 17:25:10 +00:00

ai-services-boilerplate

Repeatable build for an AI agent services host: one VM running Forgejo, Supabase and SurrealDB, joined to a Tailnet, with per-agent SSH keys, API tokens and git worktrees so several AI agents can work in parallel without touching each other.

Built from lessons measured on this fleet, not from defaults. See docs/DESIGN-NOTES.md for why each choice was made — several of them contradict what the popular templates do.


The three layers, and why they are separate

Layer Tool Runs Re-runnable
Provision the VM OpenTofu before the machine exists yes (state file)
Identity on first boot cloud-init inside the VM, once no
Everything else Ansible over SSH, any time yes, by design

cloud-init is deliberately minimal — user, keys, hostname, network, agent, and Tailscale join. Nothing else. Rented providers (DigitalOcean confirmed) freeze user-data at creation and will not let you change it, so anything you put there is a decision you cannot revise without rebuilding. Everything revisable lives in Ansible.


Quick start

# 1. provision the VM
cp terraform/terraform.tfvars.example terraform/terraform.tfvars
$EDITOR terraform/terraform.tfvars          # node, storage, bridge, IP, TS auth key
cd terraform && tofu init && tofu apply     # creates + boots the VM

# 2. configure it (run FROM the ansible dir — ansible.cfg is only read from cwd)
cd ../ansible
ansible-galaxy collection install -r requirements.yml
ansible-playbook site.yml

# 3. give each agent its own identity
cd .. && ./scripts/provision-agents.sh agent-a agent-b agent-c agent-d

Roughly 15 minutes end to end, most of it the cloud image and package installs.

Secrets are generated on the host by the playbook (.env, mode 0600) and never committed. The playbook skips regeneration if .env already exists, so re-running it does not rotate credentials out from under running containers.


What you get

  • Forgejo — git, per-agent accounts, per-agent API tokens
  • Supabase — Postgres + REST + auth, one schema per agent
  • SurrealDB — one namespace per agent
  • Caddy — the only container on the edge network, terminates TLS
  • Tailscale — unattended join via auth key, no browser click
  • Per-agent worktrees — one bare clone, N git worktrees, so agents share object storage but cannot see each other's branches

Network shape

                tailnet / LAN
                 │           │
            :80/:443       :2222 (git ssh)
                 │           │
            ┌────▼────┐      │      `edge` and `git_ssh` are the ONLY
            │  caddy  │      │      networks with a route off the box
            └────┬────┘      │
      ┌──────────┼───────┐   │
   ┌──▼───┐  ┌───▼────┐ ┌▼───────┐
   │forgejo◄──┤supabase│ │surreal │  each app on its own backend network,
   └──┬───┘  │  rest  │ │  db    │  every one marked `internal: true`
      │      └───┬────┘ └────────┘
   ┌──▼───┐  ┌───▼────┐
   │  db  │  │   db   │             ← NO route off the box, at all
   └──────┘  └────────┘

internal: true is the point. A survey of 134 templates in ChristianLempa/boilerplates-library found zero occurrences of it — backend networks there are plain bridges with full outbound NAT. That is the single line that stops a compromised container reaching the internet or your LAN.

The trap: a container on internal networks only cannot publish a port, and Docker never says so — the mapping is silently discarded while the config validates and the container reports healthy. Verified on 2026-09-16. That is why forgejo has one extra non-internal network (git_ssh) purely for :2222. Its database does not, which is the containment that actually matters. Details in docs/DESIGN-NOTES.md §5.


Layout

terraform/
  main.tf                     VM via the bpg/proxmox provider
  variables.tf                every knob, with defaults
  terraform.tfvars.example    copy to terraform.tfvars and fill in
  .terraform.lock.hcl         pins the provider version + checksum (committed)
cloud-init/
  user-data.yaml.tftpl        minimal first-boot identity; rendered by OpenTofu
                              so the Tailscale key is never written to git
ansible/
  site.yml                    docker, hardening, stack, agents, verification
  inventory.ini               addresses the host by TAILNET name
  requirements.yml            community.general + community.docker
  ansible.cfg                 pipelining + ControlPersist (only read from cwd)
compose/
  compose.yaml                the stack — internal networks, no inlined secrets
  Caddyfile                   the single ingress
  .env.example                what the playbook generates on the host
scripts/
  provision-agents.sh         per-agent keys, Forgejo accounts, tokens, worktrees
                              (DRY_RUN=1 shows the plan; credentials are masked)
  verify-agents.sh            read-only proof that each agent's token, SSH key
                              and branch actually work
  rotate-admin-token.sh       replace the Forgejo admin token and REVOKE the
                              old one (minting a new one alone leaves the old
                              one valid)
  mirror-github-user.sh       bare-mirror every repo a GitHub user owns onto
                              disk, then push into Forgejo. `sync` is the same
                              command for the first run and every update
  verify-mirror.sh            compare every branch SHA local vs Forgejo — proof the
                              copy matches, not just that the push exited 0
  gen-mirror-infopage.py      build the mirror's info page from live GitHub
                              metadata (nothing hand-typed)
  publish-org-profile.sh      put a markdown page on a Forgejo org's landing
                              page (`.profile` repo, README.md at the ROOT)
docs/
  DESIGN-NOTES.md             every decision, with the evidence behind it

Mirroring a GitHub account into the forge

Keeps a local original plus a browsable copy in Forgejo, so the agents can read a body of code without depending on GitHub being up — or on it still being there.

sudo mirror-github-user.sh sync   coleam00   # first run AND every update
sudo mirror-github-user.sh status coleam00   # what is tracked, and when
sudo verify-mirror.sh             coleam00   # prove the copy matches

sync re-reads the account each time, so new repos appear automatically — there is no list to maintain. It is idempotent: a re-run fetches only new objects and pushes only changed refs. Updates are on demand, never scheduled — nothing is in cron.

Where What
Original /srv/mirrors/github/<user>/<repo>.git clone --mirror: every branch, tag and GitHub's refs/pull/* PR history
Copy Forgejo org <user> branches + tags, browsable and clonable by the agents

Three things this gets right that a naive clone --mirror + push --mirror does not, each of which cost a failed run to find (see docs/DESIGN-NOTES.md §12):

  1. push --mirror fails on any repo that has ever had a pull request — Forgejo rejects writes to refs/pull/* because that namespace is its own. Push explicit refs/heads/* and refs/tags/* refspecs with --prune.
  2. A push does not set the default branch. Forgejo picks its own, so the forge can land visitors on the wrong branch while the local mirror is correct. It needs an explicit PATCH.
  3. Forgejo's org profile page reads README.md from the ROOT of the .profile repo — not profile/README.md, which is GitHub's convention and is silently ignored.

The control dashboard

http://10.10.10.60/control/ — paste a GitHub URL, tick the repos you want, press the button.

Paste any of: an org repositories tab URL (https://github.com/orgs/agent0ai/repositories), a profile URL, a single-repo URL, an SSH remote, or a bare owner / owner/repo. It lists every repo with stars, language, licence, size, last update and current mirror state, so you can see at a glance what is already synced. Filter, "select not yet mirrored", then run the pipeline and watch the live log.

It runs on the host under systemd, not in a container, because the work is host work (/srv/mirrors, the scripts in /usr/local/sbin) — containerising it would have meant either mounting the docker socket or duplicating the pipeline in an image.

  • Dependency-free: Python standard library only. No FastAPI, no venv, no pip. Nothing here can break on an unattended apt upgrade.
  • Unprivileged: the web service runs as aictl. Exactly one thing is escalated — mirror-pipeline.sh, by absolute path, via a single sudoers rule.
  • Not directly reachable: binds 0.0.0.0 but ufw allows 8787 only from 172.16.0.0/12 so only the Caddy container gets in. Verified refused from the tailnet.
  • One job at a time: two concurrent syncs of one account would race on the same bare repos and the same manifest.
  • The CLI and the UI run the same code — both call mirror-pipeline.sh.

The control token is generated once into /etc/ai-control/config.env (0640 root:aictl) and never rewritten, so re-running the playbook does not invalidate the token saved in your browser.

Optional keys in that same file:

Key Effect if unset
GITHUB_TOKEN Works, but GitHub rate-limits anonymous calls to 60/hour
OPENROUTER_API_KEY Metadata still shows; the per-repo Summarise button returns a clear error instead of spending money

Summaries are per click, never batched — spend tracks what you actually ask for, and the token counts are shown under each result.

Currently mirrored: coleam00 (Cole Medin) 73 repos · 1.2 GB, and agent0ai (Agent Zero) at http://10.10.10.60/agent0ai.


AI code review, keyed by BLAKE3

sudo ai-review-repo.py agent0ai agent-zero --dry-run              # free
sudo ai-review-repo.py agent0ai agent-zero --budget-usd 0.50

…or the Review button on any mirrored repo in the dashboard.

Produces one review per source file — purpose, a findings table with line numbers and severities, and notes — plus an index. Output goes to <owner>/_ai-reviews, a separate Forgejo repo, laid out as <repo>/README.md and <repo>/files/<path>.md.

Why a separate repo, not the mirror: mirrors are force-pushed on every sync, so anything committed onto a mirrored branch would be silently erased the next time you press Update.

BLAKE3 is the cache key. Every candidate file is hashed with b3sum (the Rust reference implementation); a file whose bytes are unchanged is skipped for free. Repeat runs over a large repo therefore cost almost nothing — measured on agent-zero: a second run skipped 8 of 8 already-reviewed files at zero cost.

Dependencies are skipped, deliberately. On agent-zero: 3,034 blobs → 1,104 reviewable, after dropping 700 vendored/build files (node_modules, vendor, dist, .venv, …), 1,185 non-source, 4 lockfiles, 39 trivial stubs and 2 files over 120 KB.

Spend control

Measured with inception/mercury-2.5: ≈ $0.0005 per file, so ~$0.50 for a 1,100-file repo.

Control Effect
--dry-run zero API calls; prints exactly what would be reviewed
--budget-usd hard stop, checked against OpenRouter's own reported usage.cost after every call
MAX_REVIEW_BUDGET_USD server-side ceiling (default $2) that clamps whatever the UI asks for
Account limit the OpenRouter key's own cap is the outer backstop
5 consecutive failures aborts rather than retrying into a wall

State is saved after every file, so an interrupted run resumes without re-paying for what it already did.

Treat the output as a reading order, not a verdict

On the very first run the model reported a critical defect in field(default_factory=list[str]), with correct line numbers and a confident explanation that it crashes on import. It does not — list[str] is a types.GenericAlias and is callable. The finding was false. In the same file its low finding (a bare except: pass swallowing stream errors) was real.

The prompt now carries an explicit calibration block; re-running that identical file at the same cost removed the false critical, dropped severities to low/nit, and labelled two findings Unverified. Every generated page also carries a standing caveat. Details in docs/COMPOUND-ENGINEERING.md.

A critical here means look here first, not this is broken.


Requirements

  • OpenTofu ≥ 1.12 (not Terraform — see design notes)
  • Ansible ≥ 2.18, plus ansible-galaxy collection install -r ansible/requirements.yml
  • A Proxmox API token, and a storage with content type snippets enabled
  • A Tailscale auth key — reusable, not ephemeral (an ephemeral node is reaped when offline; this is a services host other machines address by name)
  • A cloud image on the target node (…-cloudimg-…), not an installer ISO — an installer ISO has no cloud-init datasource, so none of the first-boot config would run

Validation status

This has been built and run end to end, not just linted. On 2026-09-16 it provisioned VM 460 (ai-services-460, Ubuntu 26.04.1 LTS) on a Proxmox 9.2.20 node and brought up the full stack:

Check Result
tofu fmt / validate clean against bpg/proxmox v0.113.1
tofu apply VM created, cloud-init completed, SSH by key
ansible-playbook site.yml 40 ok, 0 failed
re-run idempotent
cloud image SHA256 matched cloud-images.ubuntu.com exactly
all 7 image tags exist in their registries
services Forgejo 13.0.5, PostgREST 12.2.3, SurrealDB, Caddy — all healthy, 0 restarts
HTTP through Caddy /healthz, /api/v1/version, /db/, /surreal/health → 200
databases cannot reach the internet; unreachable from edge
docker socket not mounted by any container
git SSH :2222 bound and reachable
4 agents each token authenticates as its own user; each SSH key clones; each on its own branch

Re-run the checks yourself any time:

sudo /usr/local/sbin/verify-agents.sh     # per-agent identity, SSH, branch

Nine real defects were found and fixed during that build — the ones that validate clean and fail silently. They are documented with their evidence in docs/DESIGN-NOTES.md §5, §9, §10 and §11; the shortlist:

  • a port published from an internal-only network is silently discarded
  • iothread is silently ignored without virtio-scsi-single
  • Forgejo boots into an open install wizard that anyone can complete
  • PostgREST's "password authentication failed" was an unset password, and fixing it needs supabase_admin because postgres is not a superuser
  • Forgejo's token endpoint is not under /admin and needs basic auth
  • creating an agent account grants it no repo access

Non-goals

  • Not a cluster. One host, one VM.
  • Not multi-tenant. Every agent is trusted-but-separated, not sandboxed from root.
  • Not a Kubernetes path. Compose is the substrate — the same decision TrueNAS reached when it dropped k3s for Docker Compose in SCALE 24.10.