No description
Find a file
2026-07-09 08:31:48 -05:00
.claude chore: enable Vercel plugin project-scoped via .claude/settings.json 2026-07-09 07:24:35 -05:00
agent feat: Slack channel via direct bot token (env-based, no Vercel Connect) 2026-07-09 08:31:48 -05:00
evals fix: subagent run_sql runs without approval gate; robust evals (5/5 green) 2026-07-04 16:32:13 -05:00
test test: natural median sandbox prompt (agent queries data itself; verified local+prod) 2026-07-09 06:42:20 -05:00
.gitattributes feat: slack channel (Connect) + production build validated + LF normalization 2026-07-04 16:36:41 -05:00
.gitignore test: correct D1 refusal assertion; prod E2E green 2026-07-04 17:07:00 -05:00
.vercelignore chore: scaffold eve-analyst (eve 0.19.0 baseline) 2026-07-04 16:18:40 -05:00
AGENTS.md chore: scaffold eve-analyst (eve 0.19.0 baseline) 2026-07-04 16:18:40 -05:00
CLAUDE.md chore: scaffold eve-analyst (eve 0.19.0 baseline) 2026-07-04 16:18:40 -05:00
package.json chore: upgrade eve 0.19.0 -> 0.22.1 (+ @ai-sdk/anthropic 4.0.10); all tests green; add dry-run + record tooling 2026-07-08 13:29:12 -05:00
pnpm-lock.yaml chore: upgrade eve 0.19.0 -> 0.22.1 (+ @ai-sdk/anthropic 4.0.10); all tests green; add dry-run + record tooling 2026-07-08 13:29:12 -05:00
pnpm-workspace.yaml chore: scaffold eve-analyst (eve 0.19.0 baseline) 2026-07-04 16:18:40 -05:00
README.md test: prod correctness/edge/network-isolation checks + README 2026-07-04 21:49:58 -05:00
SPEC.md chore: scaffold eve-analyst (eve 0.19.0 baseline) 2026-07-04 16:18:40 -05:00
tsconfig.json chore: scaffold eve-analyst (eve 0.19.0 baseline) 2026-07-04 16:18:40 -05:00

eve-analyst

A production-shaped data analyst agent built on Vercel's Eve framework. You ask questions about a company's data in plain English; the agent writes read-only SQL, follows the business's own revenue rules, guards expensive queries behind human approval, runs deeper analysis in an isolated sandbox, and hands open-ended investigations to a subagent. It runs over HTTP and Slack, and its behavior is locked in by an evals suite.

Built from scratch as the companion project for a video on Eve. It mirrors Vercel's own flagship internal agent d0 (a Slack data analyst answering 30k+ questions a month), but is fully self-contained: the dataset is seeded in memory, so it runs offline and deploys in one command.

What it demonstrates (every core Eve primitive)

Eve primitive Where
An agent is a directory the whole agent/ folder
Model config agent/agent.ts (defineAgent)
System prompt in markdown agent/instructions.md
Typed tools agent/tools/ (list_tables, describe_table, run_sql, run_analysis)
Human-in-the-loop approval run_sql pauses on an unbounded full-table scan
Skills loaded on demand agent/skills/revenue-rules.md
Subagents agent/subagents/investigator/
Sandbox (isolated code) run_analysis runs Python in Vercel Sandbox (Docker locally)
Channels agent/channels/eve.ts (HTTP) + agent/channels/slack.ts
Durable sessions pause-for-approval then resume is a durable workflow
Evals as a deploy gate evals/*.eval.ts

The dataset

A tiny e-commerce warehouse (customers, products, orders, order_items, refunds) seeded deterministically into an in-memory SQLite database via node:sqlite (built into Node 24, zero dependencies). It includes test/internal accounts and refunds, so the revenue-rules skill actually changes the answer. See agent/lib/db.ts.

Requirements

  • Node 24+ (Eve requires it) and pnpm 10+
  • An ANTHROPIC_API_KEY (the agent uses @ai-sdk/anthropic with claude-sonnet-5)

Run it locally

pnpm install
export ANTHROPIC_API_KEY=sk-ant-...      # or put it in .env.local
pnpm dev                                 # starts the eve dev TUI

Then talk to it in the TUI, or drive it over HTTP:

# create a session
curl -X POST http://127.0.0.1:2000/eve/v1/session \
  -H 'content-type: application/json' \
  -d '{"message":"What was our total revenue, broken down by product category?"}'
# stream the reply (use the sessionId from the response)
curl -N http://127.0.0.1:2000/eve/v1/session/<sessionId>/stream

Try these to see each feature:

  • "What tables are in the warehouse?" → schema discovery
  • "What was our total revenue?" → loads the revenue-rules skill, answers net of refunds
  • "Run this exact query: SELECT * FROM order_items" → pauses for your approval, then resumes
  • "Use run_analysis to chart weekly revenue" → runs Python in the sandbox
  • "Revenue dropped last week, why?" → delegates to the investigator subagent

Test it

node --test test/*.test.ts             # unit tests (SQL guard + dataset)
pnpm exec eve eval --strict            # agent evals (deploy gate)
node test/e2e.mjs                      # end-to-end multi-turn conversation (needs `pnpm dev` running)
node test/prod-checks.mjs              # correctness + edge cases + sandbox network isolation (defaults to prod URL)
pnpm typecheck                         # tsc

Deploy it

vercel login          # one-time
pnpm exec eve deploy  # links a Vercel project and deploys to Vercel Functions

Set ANTHROPIC_API_KEY in the Vercel project's Environment Variables so the deployed agent can reach the model.

Enable Slack (optional)

The Slack channel (agent/channels/slack.ts) is authored and ready; credentials flow through Vercel Connect (no bot token in code). To activate:

export FF_CONNECT_ENABLED=1
vercel connect create slack --triggers
vercel connect attach <uid> --triggers --trigger-path /eve/v1/slack --yes
VERCEL_USE_EXPERIMENTAL_FRAMEWORKS=1 vercel deploy --prod

Design notes

  • Read-only by construction. agent/lib/sql-guard.ts rejects anything that is not a single SELECT; both the primary agent and the subagent go through it (agent/lib/run-select.ts).
  • Approval lives on the human-facing agent, not the subagent. The primary run_sql pauses on expensive scans; the autonomous investigator runs read-only SQL without prompts.
  • No filesystem writes. The dataset is in-memory and seeded per process, so it behaves the same on a laptop and on Vercel's read-only serverless filesystem.
  • The sandbox has no network. agent/sandbox.ts pins networkPolicy: "deny-all" (Eve's default is allow-all), so model-written code runs isolated with no egress. Verified by test/prod-checks.mjs.

License

MIT (this demo). Eve itself is Apache-2.0.