A small experiment repository comparing a base reasoning model against RLVR-GRPO checkpoints on the Math500 dataset. It includes evaluation results, short-form observations, and a local temp_clone of the full open-posttraining-system codebase for reference.
reinforcement-learningpost-trainingevaluating-modelspolicy-optimizationsparse-rewardsreasoning-modelsrlvr-grpomath500grpo-checkpointopen-posttraining-system
-
Updated
Jun 17, 2026 - Jupyter Notebook