Phoenix

AI Observability & Evaluation

Last verified:

Visit Phoenix

What is Phoenix?

Arize Phoenix is an open-source AI observability and evaluation platform for understanding, debugging, and improving AI applications. It is built on OpenTelemetry and OpenInference instrumentation, and it focuses on tracing, evaluations, prompt iteration, datasets, and experiments.

The core workflow starts with tracing, which captures what happened during a run, including model calls, retrieval, tool use, and custom logic. Phoenix then lets you attach evaluation signals to those traces so you can score outputs, spot failures, and detect regressions with more consistency than manual review.

Phoenix also supports prompt engineering workflows through captured real production examples, grouped datasets, and a prompt playground for iterating on variants. Experiments let you compare different versions of prompts, models, retrieval, or app logic on the same inputs so you can verify whether a change truly helps.

It is designed for developers and teams building LLM applications, agents, and retrieval systems who want structured debugging and measurable quality improvement. Phoenix is language-, framework-, and vendor-agnostic, with support for popular stacks like Python, TypeScript, LangChain, LlamaIndex, DSPy, Vercel AI SDK, OpenAI, Anthropic, Bedrock, and more.

Phoenix pricing

Pricing model: Freemium

Phoenix self-hosted is free and open source, with no license fees, no usage limits, and no feature gates. The website also offers Phoenix Cloud, which is presented as free to start and includes hosted instances; one page states Phoenix Developer Edition is free up to stated limits, then $50/month, while another hosted-promo page says Hosted Phoenix is free-to-use with storage for 10,000 traces and may be increased to 100,000 traces per month through signup and GitHub starring. The pricing page also lists a support add-on for self-hosted Phoenix, but the core self-hosted product itself is free. In short: self-hosted Phoenix is free; hosted/cloud Phoenix has a free tier with limits and a paid continuation path.

Phoenix pros

  • Open-source observability platform
  • OpenTelemetry-based tracing
  • OpenInference instrumentation support
  • Trace spans for step-by-step debugging
  • Covers model calls and retrieval
  • Captures tool-use execution details
  • Supports evaluation tests and scoring
  • Helps detect failures and regressions
  • Dataset creation from real runs
  • Versioned datasets for experimentation
  • Experiment comparison on identical inputs
  • Prompt playground for iteration
  • Prompt hub for saving and reusing prompts
  • Works across multiple languages
  • Works across multiple frameworks
  • Vendor-agnostic design
  • Self-hosting support
  • Cloud-hosted Phoenix option
  • Cost tracking for LLM runs
  • Built-in model pricing table
  • Custom model pricing overrides
  • Project-level analytics
  • Session-level cost analysis
  • REST API for automation

Phoenix cons

  • Primarily focused on LLM applications
  • Requires instrumentation setup
  • Best value comes after integrating traces
  • Self-hosted deployments require infrastructure management
  • Some advanced features depend on model/token metadata
  • Cost tracking needs token counts in spans
  • Custom pricing setup can be manual
  • UI can be slower in air-gapped environments if external resources stay enabled
  • Developer Edition has usage limits on the hosted version
  • Certain features are still marked as coming soon

Frequently asked questions about Phoenix

What is Arize Phoenix used for?

Arize Phoenix is used to observe, evaluate, and improve AI applications, especially LLM apps and agents. It gives you traces to see what happened during execution, evaluations to score output quality, datasets to organize examples, and experiments to compare changes on the same inputs. The goal is to move from one-off debugging to a repeatable workflow for improving quality with evidence.

What are the main features in Phoenix?

The main features are tracing, evaluation, prompt engineering, and datasets with experiments. Tracing shows model calls, retrieval, tools, and custom logic in a run. Evaluations attach quality signals to those runs, prompt tooling helps you iterate on real examples, and experiments let you compare versions systematically.

Does Phoenix support self-hosting?

Yes. Phoenix can be self-hosted and the documentation says it is free to self-host with no feature limitations. The docs also note that self-hosted deployments can run locally, in Docker, or on Kubernetes, and that data stays in your infrastructure.

Is Phoenix cloud available?

Yes. The website describes Phoenix Cloud and Hosted Phoenix as an easier way to start without managing infrastructure. It is presented as free to start, with hosted storage limits described on the site and a paid path mentioned for Phoenix Developer Edition after the free allowance.

Which frameworks and providers does Phoenix support?

Phoenix is designed to be vendor-, language-, and framework-agnostic. The docs specifically mention support for frameworks such as LangChain, LlamaIndex, DSPy, Mastra, and the Vercel AI SDK, and providers including OpenAI, Anthropic, and Bedrock, with support extending across Python, TypeScript, and Java.

How does Phoenix tracing work?

Phoenix tracing records a single application run and breaks it into spans so you can see what happened at each step. It accepts traces over OpenTelemetry and provides auto-instrumentation for popular frameworks and providers, which makes it easier to capture model calls, retrieval, tool usage, and custom code paths.

Can Phoenix track costs?

Yes. Phoenix can track token-based costs automatically from token counts and model pricing data, then roll those costs up to the span, trace, session, project, and experiment levels. It also has a built-in model pricing table and supports custom model pricing in the UI.

How do datasets and experiments help?

Datasets let you collect and version examples, including failures and edge cases, so you can reuse the same inputs for testing. Experiments then run those datasets through different versions of prompts, models, or app logic using the same evaluation criteria, making changes easier to compare fairly.

Does Phoenix have an API or SDK?

Yes. Phoenix provides a REST API for scripting workflows such as creating datasets, running experiments, querying spans, and managing prompts and projects. It also offers SDK packages, including Python client, OTEL, and evals packages, plus JavaScript support for integrations and instrumentation.

Who is Phoenix for?

Phoenix is aimed at developers, AI engineers, and teams building LLM-powered applications, agents, and retrieval workflows. It is especially useful for people who want observability, evaluation, prompt iteration, and controlled experimentation in one open-source platform.

Categories

Use cases

Browse all AI tools on NeedAnAI