Arize Phoenix
Open-source LLM evaluation and observability
Overview
Phoenix is Arize AI's open-source observability and evaluation tool for LLM applications — traces, evals, drift monitoring. Runs locally or in production; integrates with the broader Arize ML observability platform.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 20 May 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
Use cases
What you can produce with Arize Phoenix
- Trace every step of an LLM app or agent run — prompts, tool calls, retrieved documents, latencies, and token counts — in a visual span-tree UI.
- Instrument a LangGraph, LlamaIndex, CrewAI, or OpenAI Agents SDK application with a few lines of OpenTelemetry-based code.
- Run LLM-as-a-judge evaluations for hallucination, relevance, and toxicity across production traces and score them at scale.
- Debug a RAG pipeline by inspecting exactly which chunks were retrieved for a query and how they influenced the answer.
- Curate failing traces into versioned datasets and replay them as regression tests when you change prompts or models.
- Compare prompt and model variants side by side in the built-in playground using real captured production inputs.
- Self-host the entire stack as a single Docker container or pip install so no trace data ever leaves your infrastructure.
ASEAN Perspective
Arize Phoenix in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Arize Phoenix is an open-source observability and evaluation toolkit for LLM and AI applications, offering tracing, prompt and RAG debugging, evals, and OpenTelemetry-based instrumentation that runs locally or self-hosted. For AI engineers it is one of the most capable free options for understanding why an LLM pipeline behaves the way it does.
It suits developers building and debugging RAG and agent systems who want vendor-neutral tracing without paying upfront; teams wanting fully managed SLAs may prefer Arize's commercial platform or rivals like LangSmith. There is a learning curve, and deep production monitoring nudges you toward the paid tier. Open-source and self-hostable, so ASEAN use and data control are straightforward.
What people say
Phoenix is Arize AI's open-source observability and evaluation tool for LLM applications, and by 2026 it has become one of the standard names developers actually run, with over 9,000 GitHub stars and adoption at companies like DoorDash, Uber, Reddit, and Booking. It remains free and fully self-hostable under a permissive license with no feature gates, serving as the open-source front door to Arize's commercial AX platform. The product covers four connected jobs: tracing what an LLM app or agent did, running evaluations against outputs, curating failure cases into datasets, and iterating on prompts in a playground.
The thing practitioners praise most is how little it demands. Phoenix installs via pip, a single Docker container, or a Helm chart, and can run inside a Jupyter notebook — a sharp contrast with Langfuse v3, which requires wiring up ClickHouse, Redis, and S3 for self-hosting. Developers in Reddit and community threads also credit its OpenTelemetry-based instrumentation for avoiding vendor lock-in, and its integrations span most of the current agent stack: LangGraph, LlamaIndex, CrewAI, DSPy, the OpenAI and Claude agent SDKs, and Vercel AI SDK. Features that competitors paywall in self-hosted tiers, like the prompt playground and LLM-as-a-judge evals, are free in Phoenix.
The criticisms are mostly about ceiling rather than quality. Its single-process design is what makes it easy to run, but comparison writeups note it is lighter on enterprise-scale ingest, analytics, and team-management features than heavier alternatives, and organizations that outgrow it are typically steered toward the paid Arize platform. The UI is functional but less polished than Langfuse's in several head-to-heads.
Phoenix fits individual developers and small AI teams who want serious tracing and evals running locally in minutes, with zero budget and no data leaving their infrastructure. Large platform teams with heavy ingest volumes should validate its scaling story or expect an eventual migration conversation.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Arize Phoenix's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Arize Phoenix unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Arize Phoenix directly →
Spotted something out of date? Suggest an update →
Arize Phoenix in the news
More in Code & Dev Tools