Curated by

S

sushmikar-512

stacklist.com/sushmikar-512

More in AI agent evaluation frameworks

See all 10 →

Evaluation of LLM Applications - Langfuse

Langfuse Evaluation provides a comprehensive framework for checking LLM application behavior through online production trace scoring and offline experimentation with datasets, experiments, and automated or manual evaluators. The documentation covers core features including annotation queues, LLM-as-a-Judge, score analytics, CI/CD experiment integration, and code evaluators to catch regressions before deployment.

View card
Built for AI agentsACO · 500 tokens

Summary

Langfuse Evaluation provides a comprehensive framework for checking LLM application behavior through online production trace scoring and offline experimentation with datasets, experiments, and automated or manual evaluators. The documentation covers core features including annotation queues, LLM-as-a-Judge, score analytics, CI/CD experiment integration, and code evaluators to catch regressions before deployment.

Tags

langfuse · llm-evaluation · observability · datasets · experiments · llm-as-judge · ai-engineering

Key entities

Langfuse (technology, 0.99) · LLM Evaluation (concept, 0.97) · LLM-as-a-Judge (concept, 0.92) · Annotation Queues (concept, 0.85) · Datasets (concept, 0.8) · Experiments (concept, 0.85) · CI/CD (concept, 0.75)

Classification

reference · language en · status final

Provenance

claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026