AI Coach Powered by Cole's Content (RAG AI Agent)
Find a file
2025-10-26 14:14:56 -05:00
.claude AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
migrations AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
PRPs AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
src AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
tests AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
.env.example AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
.gitattributes Initial commit 2025-10-26 07:29:56 -05:00
.gitignore AI Layer in Place (Ready to Implement) 2025-10-26 10:10:34 -05:00
CLAUDE.md AI Layer in Place (Ready to Implement) 2025-10-26 10:10:34 -05:00
LICENSE Initial commit 2025-10-26 07:29:56 -05:00
pyproject.toml AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
README.md AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
SETUP_GUIDE.md AI Coach Implementation POC 2025-10-26 14:14:56 -05:00
uv.lock AI Coach Implementation POC 2025-10-26 14:14:56 -05:00

Dynamous AI Coach

RAG-powered AI coaching assistant with YouTube transcript processing pipeline.

NOTE: This code isn't fully human vetted yet since it was created as a part of my livestream. I will be refining this heavily soon!

Features

YouTube RAG Pipeline

  • Automatic transcript processing: Fetch, chunk, and index video transcripts
  • Token-aware chunking: Intelligent transcript segmentation (400-1000 tokens)
  • Vector search: Semantic search powered by Supabase + pgvector
  • Flexible embedding providers: OpenAI, Ollama, or OpenRouter
  • Timestamp preservation: Navigate directly to relevant video sections

AI Coach Agent

  • Pydantic AI agent: Supportive coaching assistant with RAG capabilities
  • Semantic search: Find relevant coaching insights across video transcripts
  • Full transcript retrieval: Get complete video transcripts with citations
  • FastAPI streaming: Real-time streaming responses via Server-Sent Events
  • JWT authentication: Secure access via Supabase Auth
  • Rate limiting: 5 requests per minute (configurable)
  • Conversation management: Auto-generated titles and message history

Quick Start

📖 See SETUP_GUIDE.md for detailed step-by-step instructions.

1. Install Dependencies

uv sync

2. Set Up Supabase Database

Run the migrations in Supabase SQL Editor:

# 1. RAG Pipeline tables (channels, videos, transcript_chunks)
# Copy contents of migrations/001_youtube_rag_schema.sql
# Paste into Supabase Dashboard > SQL Editor > Run

# 2. AI Agent tables (user_profiles, conversations, messages, requests)
# Copy contents of migrations/002_agent_tables.sql
# Paste into Supabase Dashboard > SQL Editor > Run

3. Configure Environment

Copy .env.example to .env and fill in your credentials:

cp .env.example .env
# Edit .env with your API keys

4. Run Pipeline

# Process videos from last 7 days
uv run python -m src.rag_pipeline.cli

# Custom parameters
uv run python -m src.rag_pipeline.cli --channel-id UCxxxxx --days-back 14

5. Run AI Coach Agent (Optional)

# Start the FastAPI server (default port 8030)
uv run uvicorn src.main:app --host 127.0.0.1 --port 8030 --reload

# Custom port
uv run uvicorn src.main:app --host 127.0.0.1 --port 8080 --reload

# Or use python -m to run
uv run python -m src.main

# For containers/production (listen on all interfaces)
uv run uvicorn src.main:app --host 0.0.0.0 --port 8030

Endpoints:

  • GET /health - Health check
  • POST /api/pydantic-agent - Streaming agent endpoint (requires JWT auth)

Project Structure

src/
├── agent/                   # AI Coach Agent core
│   ├── config.py           # Model & environment config
│   ├── deps.py             # Runtime dependencies
│   └── agent.py            # Agent definition with system prompt
├── tools/                   # Agent tools
│   └── rag_tools/          # RAG search and retrieval
│       ├── service.py      # Tool implementation + helpers
│       └── tool.py         # Agent tool decorators
├── api/                     # FastAPI application
│   ├── main.py             # Streaming endpoint, auth, rate limiting
│   └── db_utils.py         # Conversation & message management
├── rag_pipeline/            # YouTube transcript pipeline
│   ├── config.py           # Configuration management
│   ├── schemas.py          # Pydantic data models
│   ├── youtube_service.py  # Supadata API client
│   ├── chunking_service.py # Token-aware chunking
│   ├── embedding_service.py # Embedding generation
│   ├── storage_service.py  # Supabase vector storage
│   ├── pipeline.py         # Main orchestration
│   └── cli.py              # Command-line interface
└── utils/                   # Shared utilities
    ├── logging.py          # Structured logging
    └── clients.py          # Client initialization

tests/
├── agent/                   # Agent config tests
├── tools/rag_tools/        # RAG tools unit tests
├── api/                     # API endpoint tests
├── rag_pipeline/           # Pipeline unit tests
└── integration/            # Integration tests

Development

Lint and Type Check

# Run linter
uv run ruff check src/

# Auto-fix
uv run ruff check --fix src/

# Type check
uv run mypy src/

Run Tests

# All tests
uv run pytest tests/ -v

# Unit tests only
uv run pytest tests/ -v -m unit

# Integration tests
uv run pytest tests/ -m integration

Architecture

This project follows the vertical slice architecture with strict type safety:

  • Each feature is a self-contained slice
  • 100% type annotations (strict mypy)
  • Google-style docstrings
  • Structured logging for AI debugging

License

MIT