WARP.md
Documents a RAG service architecture, development commands, and component extension patterns for a medical device troubleshooting tool.
What this file does
Documents a RAG service architecture, development commands, and component extension patterns for a medical device troubleshooting tool.
When to use it
- Onboarding new developers to a RAG project with multiple LLM providers
- Standardizing development workflows and code quality checks across a team
- Documenting how to add new vector stores or LLM providers via abstract base classes
- Explaining hallucination prevention strategy and retrieval pipeline design
Assumes this stack
WARP.md
This file provides guidance to WARP (warp.dev) when working with code in this repository.
Project Overview
RAG Troubleshooter is a production-ready RAG service for medical device troubleshooting with vendor-neutral LLM support. It provides hallucination-safe, explainable troubleshooting guidance by ingesting service manual PDFs into a vector database and generating grounded answers via Claude, OpenAI, or Gemini.
Development Commands
Environment Setup
# Install dependencies
poetry install
# Or with pip
pip install -e .
# Configure environment variables
Copy-Item .env.example .env
# Then edit .env with your API keys
Running the Server
# Run with Poetry (recommended for development with hot reload)
poetry run python app/main.py
# Or with uvicorn directly
poetry run uvicorn app.main:app --reload
# Production mode (no reload)
poetry run uvicorn app.main:app --host 0.0.0.0 --port 8000
Server endpoints:
- Demo UI: http://localhost:8000
- API Docs: http://localhost:8000/docs
- Health Check: http://localhost:8000/api/
Document Ingestion
# Run the ingestion script to load device manuals
poetry run python ingest_devices.py
# Or via API (requires running server)
curl -X POST http://localhost:8000/api/ingest `
-H "Content-Type: application/json" `
-H "X-API-Key: dev-key-123" `
-d '{"device_name": "Bair Hugger Model 775", "filename": "bair_hugger_manual.pdf"}'
Testing
# Run tests (when implemented)
poetry run pytest
# Run tests with coverage
poetry run pytest --cov=app --cov-report=html
Code Quality
# Format code (line length: 100)
poetry run black app/
# Lint code
poetry run ruff app/
# Type check
poetry run mypy app/
Architecture Overview
Core Design Principles
- Pluggable Components: Abstract base classes enable easy swapping of vector stores and LLM providers
- Hallucination Prevention: Strict prompts, low temperature (0.1), required citations, "No grounded steps found" fallback
- Async Throughout: All I/O operations are async for scalability
- Stateless API: Easy horizontal scaling
Key Architectural Patterns
1. Provider Abstraction Pattern
Both vector stores and LLM providers use abstract base classes to enable pluggability:
-
Vector Store:
VectorStore(base) →LocalVectorStore(FAISS implementation)- Easy to add
PineconeVectorStore,WeaviateVectorStore, etc. - All implement:
add_documents(),search(),delete_by_metadata(),get_stats()
- Easy to add
-
LLM Providers:
LLMProvider(base) →ClaudeProvider,OpenAIProvider,GeminiProvider- Unified interface:
generate()method returns standardizedLLMResponse - Configuration via
LLMConfigdataclass
- Unified interface:
2. Service Layer Architecture
Three main services coordinate the RAG pipeline:
-
IngestionService: PDF → chunks → embeddings → vector store
- Handles PDF parsing, text chunking with overlap, batch embedding generation
- Chunking is sentence-boundary aware (breaks at
.or\nwhen possible)
-
RetrievalService: Query → embeddings → semantic search → re-ranking → context
- Two-stage retrieval: semantic search (FAISS) + lightweight keyword re-ranking
- Re-ranking formula:
0.7 * semantic_score + 0.2 * keyword_overlap + 0.1 * term_density
-
RAGService: Query + context → LLM → grounded answer
- Enforces hallucination-averse system prompt
- Supports multi-provider comparison (same context, different LLMs)
3. Hallucination Prevention Strategy
The system is designed to prevent hallucinations through multiple mechanisms:
- System Prompt: Explicitly instructs models to use ONLY provided context or respond with "No grounded steps found"
- Low Temperature: Set to 0.1 (configurable) to minimize creativity
- Token Limit: 500 tokens max (configurable) forces concise answers
- Citation Requirement: Every answer must cite sources with
[Source N]notation - Context Validation: Returns early with "No grounded steps found" if retrieval yields no results
4. Retrieval Pipeline
The retrieval process follows this flow:
-
Chunking (Ingestion):
- 512-character chunks with 50-character overlap (configurable)
- Sentence-boundary aware splitting to avoid mid-sentence breaks
- Metadata attached: device, page, chunk_index, source_file
-
Semantic Search:
- Query embedded using
all-MiniLM-L6-v2(configurable) - FAISS cosine similarity search
- Optional device-based metadata filtering
- Retrieves 2x requested chunks for re-ranking
- Query embedded using
-
Re-ranking:
- Combines semantic score with keyword overlap and term density
- Returns top-k chunks with highest combined scores
-
Context Formatting:
- Formats retrieved chunks as:
[Source N] (Page X, Device):\n{text} - Builds citation list with chunk IDs, pages, relevance scores
- Formats retrieved chunks as:
Data Flow
Query → RetrievalService → VectorStore → [chunks]
→ RAGService → LLMProvider → Answer + Citations
For multi-provider comparison:
Query → RetrievalService (once) → [shared context]
→ RAGService → [Claude, OpenAI, Gemini] → [Answer1, Answer2, Answer3]
Configuration System
Configuration uses Pydantic Settings with .env file support:
- API Keys:
ANTHROPIC_API_KEY,OPENAI_API_KEY,GOOGLE_API_KEY - Vector Store:
VECTOR_STORE_TYPE,VECTOR_STORE_PATH - Embeddings:
EMBEDDING_MODEL - LLM:
DEFAULT_LLM_PROVIDER,LLM_TEMPERATURE,LLM_MAX_TOKENS - Retrieval:
TOP_K_CHUNKS,CHUNK_SIZE,CHUNK_OVERLAP
All settings are centralized in app/config.py and loaded via Settings class.
Telemetry & Logging
Structured JSON logging to ./logs/rag_YYYYMMDD.log tracks:
- Query Events: Query, device, provider, model, latency, tokens, citations, has_context
- Comparison Events: Query, device, providers, latencies, tokens
- Ingestion Events: Device, source_file, pages, chunks, duration
- Error Events: Operation, error, details
Access via global telemetry instance from app.telemetry.
File Organization
app/
├── api/
│ └── routes.py # FastAPI endpoints, request/response models
├── llm_providers/
│ ├── base.py # LLMProvider abstract base class
│ ├── claude.py # Anthropic Claude implementation
│ ├── openai_provider.py # OpenAI implementation
│ └── gemini.py # Google Gemini implementation
├── services/
│ ├── ingestion.py # PDF → chunks → embeddings → vector store
│ ├── retrieval.py # Query → semantic search → re-ranking
│ └── rag.py # Retrieval + LLM generation orchestration
├── vector_store/
│ ├── base.py # VectorStore abstract base class
│ └── local_store.py # FAISS-based local implementation
├── ui/ # Static HTML demo UI
├── config.py # Pydantic Settings for env-based config
├── telemetry.py # Structured JSON logging
└── main.py # FastAPI app initialization
Adding New Components
Adding a New Vector Store
-
Create
app/vector_store/your_store.py:from app.vector_store.base import VectorStore, VectorStoreConfig class YourVectorStore(VectorStore): async def add_documents(self, texts, embeddings, metadata): # Implementation pass async def search(self, query_embedding, top_k, filter_metadata): # Implementation pass # Implement other abstract methods -
Update
app/api/routes.py→get_vector_store()to instantiate your store based on config -
Update
.env:VECTOR_STORE_TYPE=your_store
Adding a New LLM Provider
-
Create
app/llm_providers/your_provider.py:from app.llm_providers.base import LLMProvider, LLMResponse class YourProvider(LLMProvider): async def generate(self, prompt, system_prompt): # Call your LLM API return LLMResponse( text=response_text, provider="your_provider", model="model_name", tokens_used=token_count, finish_reason="stop" ) def get_provider_name(self): return "your_provider" -
Update
app/api/routes.py→get_llm_provider()to handle your provider name -
Add API key to
.env:YOUR_PROVIDER_API_KEY=...
Important Implementation Details
Embedding Model
The system uses all-MiniLM-L6-v2 by default (384 dimensions). If changing models:
- Update
EMBEDDING_MODELin.env - Update
embedding_dimensioninVectorStoreConfig(if dimensions differ) - Re-ingest all documents (embeddings are not compatible across models)
Authentication
API endpoints (except health check) require X-API-Key header matching API_KEY from .env. Default for dev: dev-key-123.
Device Management
Devices are identified by name strings (e.g., "Bair Hugger Model 775"). Device filtering in queries:
- Uses metadata filtering in vector store search
- Exact match on
devicemetadata field - Device names are case-sensitive
Context Window
The system retrieves TOP_K_CHUNKS (default: 5) chunks. Each chunk is ~512 characters. This provides ~2,500 characters of context to the LLM, well within all modern LLM context limits.
Troubleshooting
No Results Returned
- Check if PDFs have been ingested:
GET /api/stats(requires API key) - Verify device name matches exactly (case-sensitive)
- Check
./logs/rag_YYYYMMDD.logfor errors
Provider Errors
- Verify API keys in
.env - Check provider-specific rate limits
- Use
/api/compareto test all providers (gracefully handles missing keys)
Vector Store Issues
- Check
./data/vector_store/exists and has write permissions - FAISS index is created on first ingestion
- To reset: delete
./data/vector_store/and re-run ingestion
API Endpoint Summary
All endpoints except health check require X-API-Key header.
GET /api/- Health check (no auth required)GET /api/stats- Vector store statistics (total chunks, devices, embedding dimension)POST /api/query- Query with single provider (body:{query, device?, provider?})POST /api/compare- Compare all available providers (body:{query, device?})POST /api/ingest- Ingest PDF fromdata/directory (body:{device_name, filename})POST /api/export- Export answer as JSON for ServiceNow/Jira integration
See http://localhost:8000/docs for interactive API documentation.
What's inside
14 sections covering project overview, commands, architecture, file layout, component guides, troubleshooting, and API endpoints.
Change this for your project
- Replace
dev-key-123with your own API key for development - Replace
Bair Hugger Model 775andbair_hugger_manual.pdfwith your device names and filenames - Replace
luay458/taia_toolrepository references with your own repo URL
Where it goes
Save in docs/ or the repository root. Gives agents and new contributors a map of the codebase.
Worth borrowing
- Abstract base classes for pluggable vector stores and LLM providers with a unified interface
- Two-stage retrieval combining semantic search with keyword re-ranking using a weighted formula
- Hallucination prevention via strict system prompts, low temperature, citation requirements, and early fallback
Related Documents
Design Document: BharatSeva AI
Describes a 10-agent AWS system that helps India's informal workers access government schemes via voice-first, serverless architecture.
OpenClaw Enterprise Transformation Plan
Transforms a single-user AI agent into a dual-mode platform supporting both viral open-source and Fortune 500 enterprise deployments through phased security, IAM, audit, multi-tenancy, and Kubernetes features.
Qwen Image and Edit: Open-sourcing and Local GGUF Generations with Lightning
Documents the Qwen-Image and Qwen-Image-Edit models, covering architecture, training, benchmarks, ComfyUI setup, and prompting techniques for local GGUF deployment.
University of Guelph Rocketry Club - Complete Tech Stack
Documents the full tech stack of a university rocketry club website with AI chatbot, member management, and project showcases.