Galileo AI
AI evaluation platform for hallucination detection and quality
Galileo provides AI evaluation tools that detect hallucinations, measure response quality, and monitor LLM application performance. It offers real-time guardrails and automated quality metrics for RAG and agent systems.
Honeycomb AI
Observability platform with AI-powered debugging for LLM apps
Honeycomb extends its observability platform to LLM applications with OpenTelemetry-based tracing, natural language querying, and AI-powered root cause analysis for understanding complex AI system behavior.
Weights & Biases Weave
Trace, evaluate, and improve your LLM applications
Weave by Weights & Biases is an LLM observability toolkit that provides tracing, evaluation, and dataset management for AI applications. It integrates with the broader W&B experiment tracking ecosystem.