Problem: Manual grading of responses does not scale and is inconsistent.
Use case: Systematically assess response quality across large test suites.
Functionality: Grader agents with configurable rubrics, batch “grade all” execution, and integration with logs for historical evaluation.
