Skip to content
View michaelnny's full-sized avatar

Block or report michaelnny

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. rl4llmrl4llmPublic

    A Research-Friendly RL Framework for LLM Post-Training

    Python 1

  2. alpha_zeroalpha_zeroPublic

    A PyTorch implementation of DeepMind's AlphaZero agent to play Go and Gomoku board games

    Python 195 41

  3. deep_rl_zoodeep_rl_zooPublic

    A collection of Deep Reinforcement Learning algorithms implemented with PyTorch to solve Atari games and classic control tasks like CartPole, LunarLander, and MountainCar.

    Python 125 14

  4. muzeromuzeroPublic

    A PyTorch implementation of DeepMind's MuZero agent

    Python 38 7

  5. InstructLLaMAInstructLLaMAPublic

    Implements pre-training, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF), to train and fine-tune the LLaMA2 model to follow human instructions, similar to Instru…

    Jupyter Notebook 57 13

  6. Llama3-FunctionCallingLlama3-FunctionCallingPublic

    Fine-tune Llama3 model to support function calling

    Jupyter Notebook 47 7