Portable, content-addressed reliability evidence for LLM systems. Capture how a model behaves under perturbation; preserve, verify, and diff the evidence across model changes.
pythonprovenanceregression-testingai-safetydifferential-testingai-evaluationllmfalsificationllm-evaluationeu-ai-actreliability-testingmodel-migrationperturbation-testingreplayable-evidenceevaluation-validityevidence-infrastructure
-
Updated
Jun 12, 2026 - Python