Evaluations
Evaluations
Rubrics, judges, and groundedness checks for shipping model changes safely.
evaluationsintermediate
Groundedness
Measure whether answer statements are entailed by supplied evidence.
evalsgroundednesscitations
3 refs
evaluationsadvanced
Hallucination Detection
Flag answers that introduce claims not supported by retrieved context.
evalshallucinationrag
3 refs
evaluationsintermediate
LLM-as-Judge
Use a rubric-driven model judge to score subjective answer quality.
evalsjudgerubric
3 refs