Curated by
More in AI agent evaluation frameworks
See all 10 →Evaluation of LLM Applications - Langfuse
Langfuse Evaluation provides a comprehensive framework for checking LLM application behavior through online production trace scoring and offline experimentation with datasets, experiments, and automated or manual evaluators. The documentation covers core features including annotation queues, LLM-as-a-Judge, score analytics, CI/CD experiment integration, and code evaluators to catch regressions before deployment.
Built for AI agentsACO · 500 tokens
Summary
Langfuse Evaluation provides a comprehensive framework for checking LLM application behavior through online production trace scoring and offline experimentation with datasets, experiments, and automated or manual evaluators. The documentation covers core features including annotation queues, LLM-as-a-Judge, score analytics, CI/CD experiment integration, and code evaluators to catch regressions before deployment.
Tags
langfuse · llm-evaluation · observability · datasets · experiments · llm-as-judge · ai-engineering
Key entities
Langfuse (technology, 0.99) · LLM Evaluation (concept, 0.97) · LLM-as-a-Judge (concept, 0.92) · Annotation Queues (concept, 0.85) · Datasets (concept, 0.8) · Experiments (concept, 0.85) · CI/CD (concept, 0.75)
Classification
reference · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026