W2S-Eval is a PyTorch framework for AI Superalignment research. It uses confidence-gated filtering so a weaker supervisor can fine-tune a stronger student model without passing on its errors or hallucinations. This allows high-capacity models to exceed their teacher's accuracy ceiling and achieve weak-to-strong generalization.
-
Updated
Aug 10, 2026 - Python