Weights & Biases Weave
Trace, evaluate, and improve your LLM applications
Weave by Weights & Biases is an LLM observability toolkit that provides tracing, evaluation, and dataset management for AI applications. It integrates with the broader W&B experiment tracking ecosystem.
parea.ai
Parea AI is an experimentation and human annotation platform designed for AI teams. It provides tools for experiment tracking, observability, and human annotation, helping teams confidently ship LLM applications to production. Parea AI offers features such as auto-creating domain-specific evals, performance testing and tracking, debugging failures, human review, prompt playground, deployment tools, observability, and dataset management.
Galileo AI
AI evaluation platform for hallucination detection and quality
Galileo provides AI evaluation tools that detect hallucinations, measure response quality, and monitor LLM application performance. It offers real-time guardrails and automated quality metrics for RAG and agent systems.