Curated by
Measure Performance with Evaluations - Phoenix
Phoenix Evaluations is a tutorial guide for setting up and running evaluations on existing trace data to measure LLM output quality in a repeatable way. It walks through defining an LLM-as-a-judge evaluation for completeness, using a financial analysis chatbot as the example application.
Built for AI agentsACO · 2537 tokens
Summary
Phoenix Evaluations is a tutorial guide for setting up and running evaluations on existing trace data to measure LLM output quality in a repeatable way. It walks through defining an LLM-as-a-judge evaluation for completeness, using a financial analysis chatbot as the example application.
Tags
phoenix · evaluations · llm-as-a-judge · tracing · model-quality · observability · ai-evaluation
Key entities
Phoenix (technology, 0.99) · LLM-as-a-judge (concept, 0.95) · evaluations (concept, 0.95) · tracing (concept, 0.85) · Financial Analysis and Research Chatbot (technology, 0.8) · completeness evaluation (concept, 0.8)
Classification
tutorial · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026