Terraform/cloud-init/Ansible for VM 460 and the Windows pair, the ai-control dashboard, the a2a-gateway, and docs/ including the compound-engineering failure ledger, the Rust-port PRD and the agent mesh PRD. Secrets are excluded by .gitignore (*.tfvars, .env, *.token); the Proxmox token lives outside the repo at ~/.config/ai-services/pve.tfvars. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| a2a-gateway | ||
| ansible | ||
| cloud-init | ||
| compose | ||
| control | ||
| docs | ||
| scripts | ||
| terraform | ||
| terraform-windows | ||
| .gitignore | ||
| README.md | ||
ai-services-boilerplate
Repeatable build for an AI agent services host: one VM running Forgejo, Supabase and SurrealDB, joined to a Tailnet, with per-agent SSH keys, API tokens and git worktrees so several AI agents can work in parallel without touching each other.
Built from lessons measured on this fleet, not from defaults. See
docs/DESIGN-NOTES.md for why each choice was made —
several of them contradict what the popular templates do.
The three layers, and why they are separate
| Layer | Tool | Runs | Re-runnable |
|---|---|---|---|
| Provision the VM | OpenTofu | before the machine exists | yes (state file) |
| Identity on first boot | cloud-init | inside the VM, once | no |
| Everything else | Ansible | over SSH, any time | yes, by design |
cloud-init is deliberately minimal — user, keys, hostname, network, agent, and Tailscale join. Nothing else. Rented providers (DigitalOcean confirmed) freeze user-data at creation and will not let you change it, so anything you put there is a decision you cannot revise without rebuilding. Everything revisable lives in Ansible.
Quick start
# 1. provision the VM
cp terraform/terraform.tfvars.example terraform/terraform.tfvars
$EDITOR terraform/terraform.tfvars # node, storage, bridge, IP, TS auth key
cd terraform && tofu init && tofu apply # creates + boots the VM
# 2. configure it (run FROM the ansible dir — ansible.cfg is only read from cwd)
cd ../ansible
ansible-galaxy collection install -r requirements.yml
ansible-playbook site.yml
# 3. give each agent its own identity
cd .. && ./scripts/provision-agents.sh agent-a agent-b agent-c agent-d
Roughly 15 minutes end to end, most of it the cloud image and package installs.
Secrets are generated on the host by the playbook (.env, mode 0600) and
never committed. The playbook skips regeneration if .env already exists, so
re-running it does not rotate credentials out from under running containers.
What you get
- Forgejo — git, per-agent accounts, per-agent API tokens
- Supabase — Postgres + REST + auth, one schema per agent
- SurrealDB — one namespace per agent
- Caddy — the only container on the edge network, terminates TLS
- Tailscale — unattended join via auth key, no browser click
- Per-agent worktrees — one bare clone, N
git worktrees, so agents share object storage but cannot see each other's branches
Network shape
tailnet / LAN
│ │
:80/:443 :2222 (git ssh)
│ │
┌────▼────┐ │ `edge` and `git_ssh` are the ONLY
│ caddy │ │ networks with a route off the box
└────┬────┘ │
┌──────────┼───────┐ │
┌──▼───┐ ┌───▼────┐ ┌▼───────┐
│forgejo◄──┤supabase│ │surreal │ each app on its own backend network,
└──┬───┘ │ rest │ │ db │ every one marked `internal: true`
│ └───┬────┘ └────────┘
┌──▼───┐ ┌───▼────┐
│ db │ │ db │ ← NO route off the box, at all
└──────┘ └────────┘
internal: true is the point. A survey of 134 templates in
ChristianLempa/boilerplates-library found zero occurrences of it — backend
networks there are plain bridges with full outbound NAT. That is the single line
that stops a compromised container reaching the internet or your LAN.
The trap: a container on internal networks only cannot publish a port,
and Docker never says so — the mapping is silently discarded while the config
validates and the container reports healthy. Verified on 2026-09-16. That is why
forgejo has one extra non-internal network (git_ssh) purely for :2222. Its
database does not, which is the containment that actually matters. Details in
docs/DESIGN-NOTES.md §5.
Layout
terraform/
main.tf VM via the bpg/proxmox provider
variables.tf every knob, with defaults
terraform.tfvars.example copy to terraform.tfvars and fill in
.terraform.lock.hcl pins the provider version + checksum (committed)
cloud-init/
user-data.yaml.tftpl minimal first-boot identity; rendered by OpenTofu
so the Tailscale key is never written to git
ansible/
site.yml docker, hardening, stack, agents, verification
inventory.ini addresses the host by TAILNET name
requirements.yml community.general + community.docker
ansible.cfg pipelining + ControlPersist (only read from cwd)
compose/
compose.yaml the stack — internal networks, no inlined secrets
Caddyfile the single ingress
.env.example what the playbook generates on the host
scripts/
provision-agents.sh per-agent keys, Forgejo accounts, tokens, worktrees
(DRY_RUN=1 shows the plan; credentials are masked)
verify-agents.sh read-only proof that each agent's token, SSH key
and branch actually work
rotate-admin-token.sh replace the Forgejo admin token and REVOKE the
old one (minting a new one alone leaves the old
one valid)
mirror-github-user.sh bare-mirror every repo a GitHub user owns onto
disk, then push into Forgejo. `sync` is the same
command for the first run and every update
verify-mirror.sh compare every branch SHA local vs Forgejo — proof the
copy matches, not just that the push exited 0
gen-mirror-infopage.py build the mirror's info page from live GitHub
metadata (nothing hand-typed)
publish-org-profile.sh put a markdown page on a Forgejo org's landing
page (`.profile` repo, README.md at the ROOT)
docs/
DESIGN-NOTES.md every decision, with the evidence behind it
Mirroring a GitHub account into the forge
Keeps a local original plus a browsable copy in Forgejo, so the agents can read a body of code without depending on GitHub being up — or on it still being there.
sudo mirror-github-user.sh sync coleam00 # first run AND every update
sudo mirror-github-user.sh status coleam00 # what is tracked, and when
sudo verify-mirror.sh coleam00 # prove the copy matches
sync re-reads the account each time, so new repos appear automatically —
there is no list to maintain. It is idempotent: a re-run fetches only new
objects and pushes only changed refs. Updates are on demand, never
scheduled — nothing is in cron.
| Where | What | |
|---|---|---|
| Original | /srv/mirrors/github/<user>/<repo>.git |
clone --mirror: every branch, tag and GitHub's refs/pull/* PR history |
| Copy | Forgejo org <user> |
branches + tags, browsable and clonable by the agents |
Three things this gets right that a naive clone --mirror + push --mirror
does not, each of which cost a failed run to find (see
docs/DESIGN-NOTES.md §12):
push --mirrorfails on any repo that has ever had a pull request — Forgejo rejects writes torefs/pull/*because that namespace is its own. Push explicitrefs/heads/*andrefs/tags/*refspecs with--prune.- A push does not set the default branch. Forgejo picks its own, so the
forge can land visitors on the wrong branch while the local mirror is
correct. It needs an explicit
PATCH. - Forgejo's org profile page reads
README.mdfrom the ROOT of the.profilerepo — notprofile/README.md, which is GitHub's convention and is silently ignored.
The control dashboard
http://10.10.10.60/control/ — paste a GitHub URL, tick the repos you want, press the button.
Paste any of: an org repositories tab URL
(https://github.com/orgs/agent0ai/repositories), a profile URL, a single-repo
URL, an SSH remote, or a bare owner / owner/repo. It lists every repo with
stars, language, licence, size, last update and current mirror state, so you
can see at a glance what is already synced. Filter, "select not yet mirrored",
then run the pipeline and watch the live log.
It runs on the host under systemd, not in a container, because the work is
host work (/srv/mirrors, the scripts in /usr/local/sbin) — containerising it
would have meant either mounting the docker socket or duplicating the pipeline
in an image.
- Dependency-free: Python standard library only. No FastAPI, no venv, no
pip. Nothing here can break on an unattended
apt upgrade. - Unprivileged: the web service runs as
aictl. Exactly one thing is escalated —mirror-pipeline.sh, by absolute path, via a single sudoers rule. - Not directly reachable: binds
0.0.0.0but ufw allows 8787 only from172.16.0.0/12so only the Caddy container gets in. Verified refused from the tailnet. - One job at a time: two concurrent syncs of one account would race on the same bare repos and the same manifest.
- The CLI and the UI run the same code — both call
mirror-pipeline.sh.
The control token is generated once into /etc/ai-control/config.env (0640
root:aictl) and never rewritten, so re-running the playbook does not
invalidate the token saved in your browser.
Optional keys in that same file:
| Key | Effect if unset |
|---|---|
GITHUB_TOKEN |
Works, but GitHub rate-limits anonymous calls to 60/hour |
OPENROUTER_API_KEY |
Metadata still shows; the per-repo Summarise button returns a clear error instead of spending money |
Summaries are per click, never batched — spend tracks what you actually ask for, and the token counts are shown under each result.
Currently mirrored: coleam00 (Cole Medin) 73 repos · 1.2 GB, and
agent0ai (Agent Zero) at http://10.10.10.60/agent0ai.
AI code review, keyed by BLAKE3
sudo ai-review-repo.py agent0ai agent-zero --dry-run # free
sudo ai-review-repo.py agent0ai agent-zero --budget-usd 0.50
…or the Review button on any mirrored repo in the dashboard.
Produces one review per source file — purpose, a findings table with line
numbers and severities, and notes — plus an index. Output goes to
<owner>/_ai-reviews, a separate Forgejo repo, laid out as
<repo>/README.md and <repo>/files/<path>.md.
Why a separate repo, not the mirror: mirrors are force-pushed on every sync, so anything committed onto a mirrored branch would be silently erased the next time you press Update.
BLAKE3 is the cache key. Every candidate file is hashed with b3sum (the
Rust reference implementation); a file whose bytes are unchanged is skipped
for free. Repeat runs over a large repo therefore cost almost nothing —
measured on agent-zero: a second run skipped 8 of 8 already-reviewed files at
zero cost.
Dependencies are skipped, deliberately. On agent-zero: 3,034 blobs →
1,104 reviewable, after dropping 700 vendored/build files
(node_modules, vendor, dist, .venv, …), 1,185 non-source, 4 lockfiles,
39 trivial stubs and 2 files over 120 KB.
Spend control
Measured with inception/mercury-2.5: ≈ $0.0005 per file, so ~$0.50 for a
1,100-file repo.
| Control | Effect |
|---|---|
--dry-run |
zero API calls; prints exactly what would be reviewed |
--budget-usd |
hard stop, checked against OpenRouter's own reported usage.cost after every call |
MAX_REVIEW_BUDGET_USD |
server-side ceiling (default $2) that clamps whatever the UI asks for |
| Account limit | the OpenRouter key's own cap is the outer backstop |
| 5 consecutive failures | aborts rather than retrying into a wall |
State is saved after every file, so an interrupted run resumes without re-paying for what it already did.
Treat the output as a reading order, not a verdict
On the very first run the model reported a critical defect in
field(default_factory=list[str]), with correct line numbers and a confident
explanation that it crashes on import. It does not — list[str] is a
types.GenericAlias and is callable. The finding was false. In the same file
its low finding (a bare except: pass swallowing stream errors) was real.
The prompt now carries an explicit calibration block; re-running that identical
file at the same cost removed the false critical, dropped severities to
low/nit, and labelled two findings Unverified. Every generated page also
carries a standing caveat. Details in
docs/COMPOUND-ENGINEERING.md.
A
criticalhere means look here first, not this is broken.
Requirements
- OpenTofu ≥ 1.12 (not Terraform — see design notes)
- Ansible ≥ 2.18, plus
ansible-galaxy collection install -r ansible/requirements.yml - A Proxmox API token, and a storage with content type
snippetsenabled - A Tailscale auth key — reusable, not ephemeral (an ephemeral node is reaped when offline; this is a services host other machines address by name)
- A cloud image on the target node (
…-cloudimg-…), not an installer ISO — an installer ISO has no cloud-init datasource, so none of the first-boot config would run
Validation status
This has been built and run end to end, not just linted. On 2026-09-16 it
provisioned VM 460 (ai-services-460, Ubuntu 26.04.1 LTS) on a Proxmox 9.2.20
node and brought up the full stack:
| Check | Result |
|---|---|
tofu fmt / validate |
clean against bpg/proxmox v0.113.1 |
tofu apply |
VM created, cloud-init completed, SSH by key |
ansible-playbook site.yml |
40 ok, 0 failed |
| re-run | idempotent |
| cloud image | SHA256 matched cloud-images.ubuntu.com exactly |
| all 7 image tags | exist in their registries |
| services | Forgejo 13.0.5, PostgREST 12.2.3, SurrealDB, Caddy — all healthy, 0 restarts |
| HTTP through Caddy | /healthz, /api/v1/version, /db/, /surreal/health → 200 |
| databases | cannot reach the internet; unreachable from edge |
| docker socket | not mounted by any container |
git SSH :2222 |
bound and reachable |
| 4 agents | each token authenticates as its own user; each SSH key clones; each on its own branch |
Re-run the checks yourself any time:
sudo /usr/local/sbin/verify-agents.sh # per-agent identity, SSH, branch
Nine real defects were found and fixed during that build — the ones that
validate clean and fail silently. They are documented with their evidence in
docs/DESIGN-NOTES.md §5, §9, §10 and §11; the
shortlist:
- a port published from an internal-only network is silently discarded
iothreadis silently ignored withoutvirtio-scsi-single- Forgejo boots into an open install wizard that anyone can complete
- PostgREST's "password authentication failed" was an unset password, and
fixing it needs
supabase_adminbecausepostgresis not a superuser - Forgejo's token endpoint is not under
/adminand needs basic auth - creating an agent account grants it no repo access
Non-goals
- Not a cluster. One host, one VM.
- Not multi-tenant. Every agent is trusted-but-separated, not sandboxed from root.
- Not a Kubernetes path. Compose is the substrate — the same decision TrueNAS reached when it dropped k3s for Docker Compose in SCALE 24.10.