Arize Phoenix
FreeOpen-source AI observability for LLM apps
About Arize Phoenix
Arize Phoenix is an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting of LLM applications. It provides OpenTelemetry-based tracing to capture runtime behavior, LLM-powered evaluation for benchmarking performance, versioned datasets for testing and fine-tuning, experiment tracking for prompt and model changes, a playground for optimizing prompts and comparing models, prompt management with version control and tagging, an AI engineering agent (PXI) for debugging and navigation, and a remote MCP server for connecting Claude Code, Cursor, and other MCP clients. The platform is vendor and language agnostic, with out-of-the-box support for popular frameworks like OpenAI Agents SDK, LangGraph, LlamaIndex, and providers such as OpenAI, Anthropic, and Google GenAI. It can run locally, in Jupyter notebooks, containerized, or in the cloud.
Key Features
Pros & Cons
- Open-source and free to use
- Vendor and language agnostic with broad framework support (LangGraph, LlamaIndex, OpenAI Agents SDK, etc.)
- Can run locally, in Jupyter, containerized, or in the cloud
- Built-in evaluation engine using LLMs for automated benchmarking
- Includes an AI engineering agent (PXI) for debugging and navigation
- Supports remote MCP server for integration with Claude Code, Cursor, and other tools
- Primarily focused on LLM observability, may not cover traditional ML model monitoring
- Requires manual setup and instrumentation for custom frameworks
- Learning curve for OpenTelemetry-based tracing configuration
Best For
Alternatives to Arize Phoenix
Langfuse
An open-source LLM engineering platform for tracing, evaluation, prompt management, and metrics. [#opensource](https://github.com/langfuse/langfuse)
Openllmetry
Open-source observability for your LLM application, based on OpenTelemetry 