Core data model, metric interfaces, and scoring engine

Artifacts using AgentEval Core (19)
Sort by:Popular

LLM-as-judge engine with OpenAI and Anthropic provider integrations
Last Release on Apr 23, 2026
Dataset loading, management, and serialization
Last Release on Apr 23, 2026
Evaluation result reporters: console, JUnit XML, and more
Last Release on Apr 23, 2026
Built-in evaluation metric implementations
Last Release on Apr 23, 2026
Embedding model integrations for OpenAI and Ollama
Last Release on Apr 23, 2026
Chaos engineering and resilience testing for AI agents
Last Release on Apr 24, 2026
Contract testing and behavioral invariant verification for AI agents
Last Release on Apr 24, 2026
Capability profiling and fingerprinting for AI agents
Last Release on Apr 24, 2026
GitHub Actions integration with Markdown reporter and PR commenting
Last Release on Apr 23, 2026
JUnit 5 extension, annotations, and assertion API for agent evaluation
Last Release on Apr 23, 2026