Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper
-
Updated
Jul 20, 2024 - Python
Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper
NeuroRVQ: Multi-Scale Biosignal Tokenization for Generative Foundation Models
A ridiculously fast Python BPE (Byte Pair Encoder) implementation written in Rust
the small distributed language model toolkit; fine-tune state-of-the-art LLMs anywhere, rapidly
This project shows how to derive the total number of training tokens from a large text dataset from 🤗 datasets with Apache Beam and Dataflow.
Python script for manipulating the existing tokenizer.
Use custom tokenizers in spacy-transformers
Package to align tokens from different tokenizations.
Use Huggingface Transformer and Tokenizers as Tensorflow Reusable SavedModels
Python library + CLI for token counting and token distribution analysis in text datasets for LLM data workflows.
DeepSeek-OCR-2-Unlimited-OCR is an advanced, experimental visual document processing and open-ended text localization dashboard. This application establishes a unified interface that allows users to swap between two premier vision-language document models: deepseek-ai/DeepSeek-OCR-2 and baidu/Unlimited-OCR.
Megatron-LM/GPT-NeoX compatible Text Encoder with 🤗Transformers AutoTokenizer.
Recreating every milestone in Machine Learning and Artificial Intelligence
ML Model designed to learn compositional structure of LEGO assemblies
Byte Pair Encoding (BPE) tokenizer tailored for the Turkish language
Compare multilingual tokenizers and models for cost, context, and deployment decisions.
Create prompts with a given token length for testing LLMs and other transformers text models.
Evaluation toolkit for Arabic tokenizers
lute V0 implements a foundational decoder-only Transformer model, inspired by modern architectures like Qwen and Llama. It uses key techniques such as Multi-Head Attention, RMS Normalization, and the SiLU activation function.
🖥️ Utilize DeepSeek-OCR-2 to effortlessly execute advanced OCR tasks, converting documents to markdown and extracting text through an intuitive web app.
Add a description, image, and links to the tokenizers topic page so that developers can more easily learn about it.
To associate your repository with the tokenizers topic, visit your repo's landing page and select "manage topics."