DeepEval
by Confident AI
Unit-tests LLM and agent outputs in a pytest-style workflow, scoring answer relevancy, faithfulness, hallucination and tool correctness in CI.
Skills
Evaluation Metric Suite
Scores outputs on faithfulness, relevancy, hallucination and bias using research-backed metric implementations.
Pytest-Style Assertions
Wraps evaluations as familiar unit tests so quality checks run in existing CI pipelines and fail loudly.
Agent Trace Evaluation
Evaluates multi-step agent traces including tool selection and task completion, not just the final answer.
Related Agents
Armature
Product analytics and evaluation for agent sessions running against MCP servers, tracking tool calls, failures, and per…
Crawl4AI
Crawls and scrapes web pages into clean Markdown for RAG pipelines, with CSS, XPath, and LLM-driven structured extracti…
dbt MCP
Exposes the dbt Semantic Layer, Discovery API, and dbt CLI over MCP so agents query governed metrics and inspect model…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…