Build a RAG Chatbot from Scratch — Part 1: Architecture and Embeddings
Series Overview
We’re building a RAG (Retrieval-Augmented Generation) chatbot that answers questions from your own documents. Think “ChatGPT for your company’s internal docs.”
Technology stack: FastAPI (Python), ChromaDB, OpenAI embeddings, GPT-4o-mini.
Architecture
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ User │────▶│ FastAPI │────▶│ ChromaDB │────▶│ Embedding│
│ Question │◀────│ Server │◀────│ (Docs) │◀────│ Model │
└──────────┘ └────┬─────┘ └──────────┘ └──────────┘
│
▼
┌──────────┐
│ GPT-4o │
│ Mini │
└──────────┘
Flow:
- Ingestion: Documents → Chunk text → Generate embeddings → Store in ChromaDB
- Query: User question → Generate question embedding → Retrieve similar chunks → Send chunks + question to LLM → Return answer
Step 1: Project Setup
pip install chromadb openai fastapi uvicorn pypdf python-dotenv
# app/config.py
from pydantic_settings import BaseSettings
class Settings(BaseSettings):
openai_api_key: str
chroma_persist_dir: str = "./chroma_data"
model: str = "gpt-4o-mini"
embedding_model: str = "text-embedding-3-small"
chunk_size: int = 500
chunk_overlap: int = 50
class Config:
env_file = ".env"
settings = Settings()
Step 2: Document Chunking
# app/ingestion.py
from langchain.text_splitter import RecursiveCharacterTextSplitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50,
separators=["\n\n", "\n", ".", " ", ""]
)
def chunk_document(text: str) -> list[str]:
"""Split document into overlapping chunks."""
return text_splitter.split_text(text)
Why overlap? Ensures sentences that cross chunk boundaries aren’t orphaned. A 500-character chunk with 50-char overlap means each sentence appears in at least 2 chunks.
Step 3: Generate Embeddings
# app/embeddings.py
from openai import OpenAI
from app.config import settings
client = OpenAI(api_key=settings.openai_api_key)
def get_embedding(text: str) -> list[float]:
"""Generate 1536-dimensional embedding for text."""
response = client.embeddings.create(
model=settings.embedding_model,
input=text
)
return response.data[0].embedding
def embed_chunks(chunks: list[str]) -> list[list[float]]:
"""Generate embeddings for a list of text chunks."""
embeddings = []
for chunk in chunks:
embeddings.append(get_embedding(chunk))
return embeddings
Cost: text-embedding-3-small costs $0.02 per 1M tokens. Embedding 1,000 pages costs approximately $0.01.
Step 4: Store in ChromaDB
# app/vector_store.py
import chromadb
from app.config import settings
chroma_client = chromadb.PersistentClient(path=settings.chroma_persist_dir)
def get_or_create_collection(name: str = "docs"):
return chroma_client.get_or_create_collection(
name=name,
metadata={"hnsw:space": "cosine"}
)
def ingest_document(doc_id: str, text: str, metadata: dict):
from app.ingestion import chunk_document
from app.embeddings import embed_chunks
collection = get_or_create_collection()
chunks = chunk_document(text)
embeddings = embed_chunks(chunks)
collection.add(
ids=[f"{doc_id}_chunk_{i}" for i in range(len(chunks))],
embeddings=embeddings,
documents=chunks,
metadatas=[{**metadata, "chunk_index": i} for i in range(len(chunks))]
)
return len(chunks)
Step 5: Ingestion API
# app/api/ingest.py
from fastapi import APIRouter, UploadFile, File
from app.vector_store import ingest_document
import pypdf
router = APIRouter(prefix="/ingest", tags=["ingestion"])
@router.post("/pdf")
async def ingest_pdf(file: UploadFile = File(...)):
reader = pypdf.PdfReader(file.file)
text = "\n".join([page.extract_text() for page in reader.pages])
chunk_count = ingest_document(
doc_id=file.filename,
text=text,
metadata={"source": file.filename}
)
return {"filename": file.filename, "chunks_created": chunk_count}
Verification
# Create a test PDF
echo "RAG combines retrieval with generation. It retrieves relevant documents and feeds them to a language model to produce grounded answers." > test.txt
uvicorn app.main:app --reload
# Ingest text (simulate a PDF)
curl -X POST http://localhost:8000/ingest/pdf \
-F "[email protected]"
# {"filename":"test.txt","chunks_created":1}
Summary
- ChromaDB stores document embeddings with cosine similarity search
- Recursive text splitting breaks documents into overlapping 500-char chunks
- text-embedding-3-small generates 1536-dimension vectors at $0.02/1M tokens
- Ingestion API accepts PDFs, extracts text, chunks, embeds, and stores
Advertisement