
Core data model, metric interfaces, and scoring engine

LLM-as-judge engine with OpenAI and Anthropic provider integrations
Last Release on Apr 23, 2026

Dataset loading, management, and serialization
Last Release on Apr 23, 2026

Evaluation result reporters: console, JUnit XML, and more
Last Release on Apr 23, 2026

Built-in evaluation metric implementations
Last Release on Apr 23, 2026

Embedding model integrations for OpenAI and Ollama
Last Release on Apr 23, 2026

Chaos engineering and resilience testing for AI agents
Last Release on Apr 24, 2026

Contract testing and behavioral invariant verification for AI agents
Last Release on Apr 24, 2026

Capability profiling and fingerprinting for AI agents
Last Release on Apr 24, 2026

GitHub Actions integration with Markdown reporter and PR commenting
Last Release on Apr 23, 2026

JUnit 5 extension, annotations, and assertion API for agent evaluation
Last Release on Apr 23, 2026