Arize Phoenix logo

Arize Phoenix

Free

Open-source AI observability for LLM apps

Agent TracingFreeFree tier
Type
Open Source
Company
Arize AI

About Arize Phoenix

Arize Phoenix is an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting of LLM applications. It provides OpenTelemetry-based tracing to capture runtime behavior, LLM-powered evaluation for benchmarking performance, versioned datasets for testing and fine-tuning, experiment tracking for prompt and model changes, a playground for optimizing prompts and comparing models, prompt management with version control and tagging, an AI engineering agent (PXI) for debugging and navigation, and a remote MCP server for connecting Claude Code, Cursor, and other MCP clients. The platform is vendor and language agnostic, with out-of-the-box support for popular frameworks like OpenAI Agents SDK, LangGraph, LlamaIndex, and providers such as OpenAI, Anthropic, and Google GenAI. It can run locally, in Jupyter notebooks, containerized, or in the cloud.

Key Features

Tracing: OpenTelemetry-based tracing of LLM application runtime
Evaluation: LLM-powered response and retrieval evaluations for benchmarking
Datasets: Versioned datasets for experimentation, evaluation, and fine-tuning
Experiments: Track and evaluate changes to prompts, LLMs, and retrieval
Playground: Optimize prompts, compare models, adjust parameters, replay traced calls
Prompt Management: Version control, tagging, and experimentation for prompts
PXI (Phoenix Intelligence): Built-in AI agent for debugging traces and navigating the product
Remote MCP Server: Connect MCP clients (Claude Code, Cursor) to query traces and experiments

Pros & Cons

Pros
  • Open-source and free to use
  • Vendor and language agnostic with broad framework support (LangGraph, LlamaIndex, OpenAI Agents SDK, etc.)
  • Can run locally, in Jupyter, containerized, or in the cloud
  • Built-in evaluation engine using LLMs for automated benchmarking
  • Includes an AI engineering agent (PXI) for debugging and navigation
  • Supports remote MCP server for integration with Claude Code, Cursor, and other tools
Cons
  • Primarily focused on LLM observability, may not cover traditional ML model monitoring
  • Requires manual setup and instrumentation for custom frameworks
  • Learning curve for OpenTelemetry-based tracing configuration

Best For

Debugging and troubleshooting LLM application behaviorBenchmarking LLM performance with retrieval and response evaluationsExperimenting with prompt variations and model comparisonsManaging prompt versions and testing changes systematicallyMonitoring and analyzing LLM traces in production or developmentIntegrating AI observability into existing MCP-compatible workflows

Alternatives to Arize Phoenix

FAQ

What is Arize Phoenix?
Arize Phoenix is an open-source AI observability platform for experimentation, evaluation, and troubleshooting of LLM applications.
Where can I run Phoenix?
Phoenix can run locally, in a Jupyter notebook, in a containerized environment, or in the cloud.
Does Phoenix support my LLM framework?
Phoenix is vendor and language agnostic, with out-of-the-box support for frameworks like OpenAI Agents SDK, LangGraph, LlamaIndex, CrewAI, and many more.