trl
Here are 103 public repositories matching this topic...
Turn any research paper into a commercialization report — 6 AI agents, TRL/MRL scoring, patent landscape, market intelligence, verified citations. DeepSeek / OpenAI / Claude.
-
Updated
Aug 13, 2026 - Python
Notus is a collection of fine-tuned LLMs using SFT, DPO, SFT+DPO, and/or any other RLHF techniques, while always keeping a data-first approach
-
Updated
Jan 15, 2024 - Python
An implementation of GRPO for Unsloth's VLMs training
-
Updated
Aug 7, 2025 - Python
Code repository dedicated to experimenting and research with tiny reasoning language model
-
Updated
Nov 24, 2025 - Python
Distill teacher chains-of-thought into a LoRA adapter via a strict boxed-answer format contract + two-phase Train→Nudge (silver-medal NVIDIA Nemotron reasoning recipe, as a tested library).
-
Updated
Jul 27, 2026 - Python
Agentic RL 零基础中文教程:24 章从概念到 GRPO 实战,含 TRL 最小可跑示例 | Beginner-friendly Agentic RL tutorial with hands-on GRPO project
-
Updated
Jun 3, 2026 - Python
simpleR1: A Simple Framework for Training R1-like Models
-
Updated
Aug 12, 2025 - Python
使用trl、peft、transformers等库,实现对huggingface上模型的微调。
-
Updated
Mar 21, 2025 - Python
Different post-training techniques for LLMs, including: SFT, DPO and Online RL
-
Updated
Sep 5, 2025 - Python
Config-driven LLM fine-tuning with safety evaluation, EU AI Act compliance, 6 alignment methods, and one-command bundled quickstart templates.
-
Updated
Jul 21, 2026 - Python
Domain-neutral LLM post-training pipeline for CPT, Fact-SFT, optional DPO, adapter merge, quality evaluation, GGUF/ONNX export, and OpenAI-compatible inference.
-
Updated
Jul 16, 2026 - Python
DPO/SafeDPO/OPAD training + eval for teaching tool-using LLMs to refuse falsely-benign MCP exploits
-
Updated
Jul 8, 2026 - Python
Minimal code to train reasoning model with reinforcement learning.
-
Updated
Aug 9, 2025 - Python
Process-supervised RL for a multi-step reasoning agent — DAPO + a learned Process Reward Model (PRM) training a Qwen3-8B Planner. A modern, from-scratch rebuild of the AgentFlow paper (ICLR 2026).
-
Updated
Jun 2, 2026 - Python
A production-ready Python script that automatically generates high-quality supervised fine-tuning (SFT) datasets for chat-based language models using Google's Gemini 2.0 Flash API. Compatible with TRL (Transformer Reinforcement Learning) and Hugging Face Transformers.
-
Updated
Jan 4, 2026 - Python
Improve this page
Add a description, image, and links to the trl topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the trl topic, visit your repo's landing page and select "manage topics."