Trust center / Benchmark

Measure the failure modes that matter.

The benchmark is a versioned evaluation harness, not a marketing percentage. Results are published only with dataset, configuration, date, sample size and known limitations.

Evaluation dimensions

Verdict quality

Per-class precision, recall and F1 for supported, contradicted, disputed and unverifiable claims.

Selective risk

Coverage-versus-error curves and performance when the engine abstains instead of guessing.

Retrieval

Evidence recall, entailment, contradiction, independence and source-class compliance.

Extraction

Atomicity, span alignment, material-claim recall, citation integrity, dates, quantities and jurisdiction.

Reproducibility controls

  • Freeze dataset and policy versions; hash inputs and result artifacts.
  • Record provider model IDs, retrieval configuration, judge configuration and run timestamp.
  • Disable result caches for benchmark runs unless cache behavior is itself under test.
  • Run serially when provider throttling or shared retrieval state could bias outcomes.
  • Preserve raw structured outputs privately and publish aggregate, non-sensitive results.

Current publication status

No current public performance claim is asserted on this page.

The recovered system contains a historical 288-item evaluation dataset and prior result artifacts, but those numbers are not presented as current until the rebuilt contract is run end-to-end under a frozen 2026 configuration and reviewed for label and dataset validity.

Required report fields

Every public report must identify the evaluation date, dataset composition, exclusions, label process, sample size, confidence intervals where meaningful, model and policy versions, provider failures, abstention rate, coverage, false-positive analysis and reproducibility instructions.