Curated by
More in AI agent evaluation frameworks
See all 10 →Inspect: Open-source Framework for LLM Evaluations
Inspect is an AI evaluation framework developed by the UK AI Security Institute and Meridian Labs, offering composable building blocks, over 200 pre-built evaluations, and support for 20+ model providers. The framework enables coding, reasoning, knowledge, and agentic task evaluations with features including sandboxing, tool calling, multi-agent primitives, and a web-based visualization tool.
Built for AI agentsACO · 2105 tokens
Summary
Inspect is an AI evaluation framework developed by the UK AI Security Institute and Meridian Labs, offering composable building blocks, over 200 pre-built evaluations, and support for 20+ model providers. The framework enables coding, reasoning, knowledge, and agentic task evaluations with features including sandboxing, tool calling, multi-agent primitives, and a web-based visualization tool.
Tags
ai-evaluation · inspect · framework · llm-benchmarks · agentic-tasks · model-testing · python
Key entities
Inspect (technology, 1) · UK AI Security Institute (organization, 0.95) · Meridian Labs (organization, 0.9) · SimpleQA (technology, 0.85) · Claude Code (technology, 0.8) · VS Code Extension (technology, 0.75) · OpenAI (organization, 0.9) · Anthropic (organization, 0.9) · Google (organization, 0.85) · Docker (technology, 0.7) · Kubernetes (technology, 0.7) · HuggingFace (technology, 0.8) · AI evaluation (concept, 0.95) · MCP tools (technology, 0.7)
Classification
framework · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026