---
title: "Evaluation of LLM Applications - Langfuse"
url: https://stacklist.com/card/a65b7ae8-5b14-4af1-9c3f-0e2b3042dfc2
source_url: "https://langfuse.com/docs/evaluation/overview"
stack: https://stacklist.com/c/technology/stack/d244b35a-f040-4bb7-96ae-187b792f699b
summary: "Langfuse Evaluation provides a comprehensive framework for checking LLM application behavior through online production trace scoring and offline experimentation with datasets, experiments, and automated or manual evaluators. The documentation covers core features including annotation queues, LLM-as-a-Judge, score analytics, CI/CD experiment integration, and code evaluators to catch regressions before deployment."
tags: "langfuse, llm-evaluation, observability, datasets, experiments, llm-as-judge, ai-engineering"
key_entities: "Langfuse (technology), LLM Evaluation (concept), LLM-as-a-Judge (concept), Annotation Queues (concept), Datasets (concept), Experiments (concept), CI/CD (concept)"
classification: "reference"
content_hash: "sha256:992e0a68abb847adac380b7b9de67e6e8576ae9e654f161d893236ec245229cd"
acp_version: "0.2"
token_counts_approximate: 500
visibility: public
agent_accessible: true
status: "final"
---

# Evaluation of LLM Applications - Langfuse

Docs Evaluation Overview Copy page Evaluation Overview Evals give you a repeatable check of your LLM application&#x27;s behavior. You replace guesswork with data, and catch regressions before you ship a change. Evaluation runs across most of the AI engineering loop : you score live traces in production, turn interesting examples into datasets, run experiments to compare changes, and judge the results with manual or automated evaluators. It happens both online , on live production traces, and offline , before you ship a change. Deploy Online Trace traces · sessions · agents · prompts Online Monitor dashboards · LLM-as-judge · feedback Offline Build datasets datasets · features-as-tests Offline Experiment prompts · models · code variants Offline Evaluate judges · custom evals · annotation 🎥 Watch this walkthrough of Langfuse Evaluation and how to use it to improve your LLM application. Getting Started Start with the Core Concepts page. It explains how evaluators, scores, datasets, and experiments fit together in Langfuse, which makes the rest of the docs much easier to navigate. Once you have that context, use the table below to find the right feature page: If you want to... Use this Langfuse feature Review and rate traces manually Annotation Queues , Scores via UI Leave open-ended notes on traces Text scores , Annotation Queues Track recurring failure categories Score configs , scores Build a reusable set of test cases Datasets Compare prompt, model, or code changes side by side Experiments via UI , Experiments via SDK Block deploys on regressions CI/CD experiments Run deterministic checks Code Evaluators Automatically score live production traces LLM-as-a-Judge , Scores via API/SDK See how scores trend over time Score Analytics , custom dashboards Already know what you&#x27;re looking for? Browse Evaluation Methods and Experiments in the sidebar. GitHub Discussions Was this page helpful? Good Bad Support Last edited Previous Troubleshooting and FAQ Next Concepts
