Skip to content
View liziniu's full-sized avatar

Highlights

  • Pro

Block or report liziniu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. ReMaxReMaxPublic

    Code for Paper (ReMax: A Simple, Efficient and Effective Reinforcement Learning Method for Aligning Large Language Models)

    Python 201 15

  2. GEMGEMPublic

    Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)

    Python 59 6

  3. verl-project/verlverl-project/verlPublic

    verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

    Python 23.1k 4.4k

  4. zyushun/Adam-minizyushun/Adam-miniPublic

    Code for Adam-mini: Use Fewer Learning Rates To Gain More https://arxiv.org/abs/2406.16793

    Python 460 19

  5. policy_optimizationpolicy_optimizationPublic

    Code for Paper (Policy Optimization in RLHF: The Impact of Out-of-preference Data)

    Python 29 6

  6. cold_start_rlcold_start_rlPublic

    Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?

    Python 20