A production-grade LLM Evaluation & Benchmarking Framework for systematic model auditing. Features parallel benchmarking, fairness/bias detection, MMLU integration, and a real-time analytics dashboard powered by React and FastAPI.
benchmarkingperformance-analysisai-safetymlopsfastapifairness-mlresearch-toolsprompt-engineeringgenerative-aillm-evaluationreact-v19model-auditingmmlu-benchmark
-
Updated
Apr 14, 2026 - Python