Evaluation dimensions
Verdict quality
Per-class precision, recall and F1 for supported, contradicted, disputed and unverifiable claims.
Selective risk
Coverage-versus-error curves and performance when the engine abstains instead of guessing.
Retrieval
Evidence recall, entailment, contradiction, independence and source-class compliance.
Extraction
Atomicity, span alignment, material-claim recall, citation integrity, dates, quantities and jurisdiction.
Reproducibility controls
- Freeze dataset and policy versions; hash inputs and result artifacts.
- Record provider model IDs, retrieval configuration, judge configuration and run timestamp.
- Disable result caches for benchmark runs unless cache behavior is itself under test.
- Run serially when provider throttling or shared retrieval state could bias outcomes.
- Preserve raw structured outputs privately and publish aggregate, non-sensitive results.
Current publication status
The recovered system contains a historical 288-item evaluation dataset and prior result artifacts, but those numbers are not presented as current until the rebuilt contract is run end-to-end under a frozen 2026 configuration and reviewed for label and dataset validity.
Required report fields
Every public report must identify the evaluation date, dataset composition, exclusions, label process, sample size, confidence intervals where meaningful, model and policy versions, provider failures, abstention rate, coverage, false-positive analysis and reproducibility instructions.