FAIR cyber risk quantification toolkits, agent-based control simulation (FAIR-CAM), threat event frequency estimator (PyPI), LLM classification validator (PyPI), Monte Carlo risk engine with IRIS benchmarks.
-
Updated
Jul 3, 2026 - Python
FAIR cyber risk quantification toolkits, agent-based control simulation (FAIR-CAM), threat event frequency estimator (PyPI), LLM classification validator (PyPI), Monte Carlo risk engine with IRIS benchmarks.
A comprehensive Python framework for evaluating LLM-extracted structured data against ground truth labels. Supports binary classification, scalar values, and list fields with detailed performance metrics, confidence-based evaluation, and statistical uncertainty quantification via non-parametric bootstrap confidence intervals.
pytest for LLM apps - Test for grounding failures, prompt injection, safety violations, and regressions
Deterministic validation firewall that verifies AI-generated proposals against ground-truth state using immutable rules. Zero dependencies. Patent pending.
A new package that helps developers integration-test AI and LLM applications by validating structured outputs. It takes a user's test scenario or prompt as input, sends it to an LLM, and uses pattern
Setting benchmarking standards for the LLM security!
Production-grade AI quality gates for LLM applications. Automating hallucination detection, semantic accuracy scoring, and adversarial safety testing for RAG & Agentic AI systems.
GSMA telecom documentation dataset creation pipeline with hard negative generation for embedding training. Features concurrent LLM validation, semantic similarity ranking, and DVC-based reproducible data processing.
MCP-native post-generation validation layer for LLM outputs
MCP server for heterogeneous AI validation — T1 structured reasoning, T2 cross-model evaluation, and deterministic checksum. Python stdlib-only.
Contract testing for LLM structured outputs. YAML contracts bind JSON Schema, semantic invariants (sums, source-text spans, cross-field rules), and CI thresholds; a measured repair ladder rescued 5,991 of 20k labeled outputs with verdicts matching ground truth 20,000/20,000. Offline, deterministic, no API keys; exit codes gate merges.
A unified suite of MCP suppressors that prevent hallucinations, enforce grounding, and stabilise agent behaviour.
GenAI Chatbot Testing Framework using Python & Pytest | Validates intent, context, prompt variations, hallucination risks, safety, and response quality scoring
Add a description, image, and links to the llm-validation topic page so that developers can more easily learn about it.
To associate your repository with the llm-validation topic, visit your repo's landing page and select "manage topics."