FAIR cyber risk quantification toolkits, agent-based control simulation (FAIR-CAM), threat event frequency estimator (PyPI), LLM classification validator (PyPI), Monte Carlo risk engine with IRIS benchmarks.
-
Updated
Jul 3, 2026 - Python
FAIR cyber risk quantification toolkits, agent-based control simulation (FAIR-CAM), threat event frequency estimator (PyPI), LLM classification validator (PyPI), Monte Carlo risk engine with IRIS benchmarks.
A comprehensive Python framework for evaluating LLM-extracted structured data against ground truth labels. Supports binary classification, scalar values, and list fields with detailed performance metrics, confidence-based evaluation, and statistical uncertainty quantification via non-parametric bootstrap confidence intervals.
pytest for LLM apps - Test for grounding failures, prompt injection, safety violations, and regressions
Evidence-backed validation framework for AI agents, Cursor skills, MCP orchestrations, and agentic repos. Structured rule catalog, verified evidence standard, and certification output. Free, open-source, MIT.
Deterministic validation firewall that verifies AI-generated proposals against ground-truth state using immutable rules. Zero dependencies. Patent pending.
A new package that helps developers integration-test AI and LLM applications by validating structured outputs. It takes a user's test scenario or prompt as input, sends it to an LLM, and uses pattern
Systematic AI evaluation framework that transforms subjective assessment into objective measurement. Reduce research time by 85% while maintaining 95%+ accuracy through multi-LLM validation.
Setting benchmarking standards for the LLM security!
Production-grade AI quality gates for LLM applications. Automating hallucination detection, semantic accuracy scoring, and adversarial safety testing for RAG & Agentic AI systems.
GSMA telecom documentation dataset creation pipeline with hard negative generation for embedding training. Features concurrent LLM validation, semantic similarity ranking, and DVC-based reproducible data processing.
Local-LLM agentic pipeline (n8n + Ollama) that turns a raw requirement into validated QA test cases — two independent validation layers (semantic + structural), zero cloud dependency, runs on a 16GB CPU-only laptop. Includes a documented hallucination failure case and honest known-limitations section.
MCP-native post-generation validation layer for LLM outputs
MCP server for heterogeneous AI validation — T1 structured reasoning, T2 cross-model evaluation, and deterministic checksum. Python stdlib-only.
Contract testing for LLM structured outputs. YAML contracts bind JSON Schema, semantic invariants (sums, source-text spans, cross-field rules), and CI thresholds; a measured repair ladder rescued 5,991 of 20k labeled outputs with verdicts matching ground truth 20,000/20,000. Offline, deterministic, no API keys; exit codes gate merges.
A unified suite of MCP suppressors that prevent hallucinations, enforce grounding, and stabilise agent behaviour.
GenAI Chatbot Testing Framework using Python & Pytest | Validates intent, context, prompt variations, hallucination risks, safety, and response quality scoring
C# SDK for the Guardrails AI API -- LLM validation, hallucination detection, PII protection, and content safety
Add a description, image, and links to the llm-validation topic page so that developers can more easily learn about it.
To associate your repository with the llm-validation topic, visit your repo's landing page and select "manage topics."