Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Transformer

Build StatusLicensePyTorchHuggingFace CompatibleStarsDownloads

A polished PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer. Designed for clarity, reproducibility, and interoperability with HuggingFace Transformers, this repository provides a robust baseline for Research and Engineering being Fully Configurable. The codebase emphasizes readable and well-documented components so you can iterate on Feed-Forward, Attention and Normalization blocks and other architectural variants with minimal friction.

Features

  • Fully Configurable architecture (layers, heads, model dimensions, dropout, etc.)
  • HuggingFace-compatible API alignment with past_key_values support for efficient generation
  • KV-Cache support for fast incremental decoding
  • Encoder-Decoder architecture support with cross-attention
  • Vision Transformer (ViT) support for image classification and feature extraction
  • Multiple Attention Variants: MHA, GQA (Grouped Query Attention), CrossAttention
  • Flexible Position Encodings: RoPE, PartialRoPE, ALiBi
  • LoRA Integration for parameter-efficient fine-tuning
  • Flash Attention support for accelerated training and inference
  • Compact and easily extensible design for rapid prototyping and research experiments
  • Clear, well-documented modules to facilitate experimentation with attention, FFNs, etc.

Download the code

git clone --depth=1 https://github.com/lof310/transformer
cd transformer

Installation

# Install dependencies
pip install -r requirements.txt
# Install on developer mode (Recommended)
pip install -e .# Install Normally
pip install .

Quick Start

Decoder-Only Model

importtorchfromtransformerimportTransformer, TransformerConfig# Configure the modelconfig=TransformerConfig(
n_layers=12,
n_heads=32,
d_model=1536,
attn_qk_norm=True,
tied_weights=False,
seq_len=1024,
max_seq_len=4096,
)
# Initialize modelmodel=Transformer(config)
# Forward PassB, N=16, 1024input_ids=torch.randint(low=0, high=config.vocab_size, size=(B, N))
output=model(input_ids=input_ids, return_states=False)

Vision Transformer (ViT) for Image Processing

importtorchfromtransformerimportTransformer, TransformerConfig# Configure ViT for image classificationconfig=TransformerConfig(
n_layers=12,
n_heads=16,
d_model=1024,
vocab_size=1000, # Number of output classespatch_size=16, # Patch size (16x16 pixels per token)img_size=224, # Input image sizein_channels=3, # RGB imagesmax_seq_len=512, # Must accommodate (img_size/patch_size)^2 + 1 CLS tokenpos_encoding="RoPE",
)
model=Transformer(config)
model.eval()
# Process images: shape (batch_size, channels, height, width)images=torch.randn(4, 3, 224, 224)
withtorch.no_grad():
output=model(images=images)
logits=output.logits# Shape: (4, 197, 1000)# Use CLS token (first position) for classificationcls_logits=logits[:, 0, :] # Shape: (4, 1000)predictions=cls_logits.argmax(dim=-1)

Incremental Decoding with KV-Cache

importtorchfromtransformerimportTransformer, TransformerConfigconfig=TransformerConfig(n_layers=6, n_heads=8, d_model=512)
model=Transformer(config)
model.eval()
# Initial promptinput_ids=torch.randint(0, config.vocab_size, (1, 10))
# First forward pass (no cache)withtorch.no_grad():
output=model(input_ids, use_cache=True)
logits=output.logitspast_key_values=output.past_key_values# Cache for next step# Incremental decoding (one token at a time)next_token_id=logits[:, -1:].argmax(dim=-1)
withtorch.no_grad():
output=model(next_token_id, past_key_values=past_key_values, use_cache=True)
new_past_key_values=output.past_key_values# Updated cache

Encoder-Decoder Model

importtorchfromtransformerimportTransformer, TransformerConfig, EncoderDecoderModel# Encoder configencoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
pos_encoding="RoPE",
)
# Decoder config (with cross-attention)decoder_config=TransformerConfig(
n_layers=6,
n_heads=8,
d_model=512,
attn_class="GQA",
n_kv_heads=4,
pos_encoding="RoPE",
add_cross_attention=True, # Enable cross-attention
)
# Create encoder-decoder modelmodel=EncoderDecoderModel(encoder_config, decoder_config)
# Forward passencoder_input=torch.randint(0, encoder_config.vocab_size, (4, 20))
decoder_input=torch.randint(0, decoder_config.vocab_size, (4, 10))
output=model(
input_ids=decoder_input,
encoder_input_ids=encoder_input,
return_dict=True
)

Applying LoRA Adapters

fromtransformerimportapply_lora_to_model# After creating your modelmodel=Transformer(config)
# Apply LoRA to specific layers (e.g., query/key/value projections)apply_lora_to_model(model, target_modules=["qkv_proj"], lora_rank=8, lora_alpha=16)
# Now only LoRA parameters are trainableforparaminmodel.parameters():
param.requires_grad=Falseforname, paraminmodel.named_parameters():
if"lora_"inname:
param.requires_grad=True

Default Configuration

The default configuration implements the latest SOTA Transformer design.

fromtransformerimportTransformerConfigTransformerConfig(
n_layers=12,
d_model=1536,
n_heads=32,
n_kv_heads=None, # GQA Disabled (MHA by default)vocab_size=50000,
d_ff=None, # Chosen Automatically (ratio 8/3 ≈ 2.666)norm_design="pre_norm",
norm_class="rms_norm",
ffn_class="SwiGLU",
attn_class="MHA", # Options: "MHA", "GQA", "CrossAttention"block_class=None, # Uses default TransformerBlockattn_bias=False,
ffn_bias=True,
lm_head_bias=False,
attn_qk_norm=True,
attn_dropout=0.0,
tied_weights=False,
seq_len=1024,
pos_encoding="RoPE", # Options: "RoPE", "PartialRoPE", "ALiBi"rope_base=10000.0,
max_seq_len=4096,
add_cross_attention=False, # Enable for encoder-decoder
)

Architecture Overview

Attention Mechanisms

  • MHA (Multi-Head Attention): Standard self-attention with equal query/key/value heads
  • GQA (Grouped Query Attention): Efficient attention with fewer KV heads, sharing across query groups
  • CrossAttention: Attention between decoder queries and encoder key/values for seq2seq tasks

Position Encodings

  • RoPE (Rotary Position Embeddings): Rotates query/key vectors based on absolute positions
  • PartialRoPE: Applies RoPE to only a subset of dimensions
  • ALiBi (Attention with Linear Biases): Adds distance-based biases to attention scores

Normalization Designs

  • pre_norm: Normalize before attention/FFN (recommended for deep models)
  • post_norm: Normalize after attention/FFN (original Transformer)
  • parallel: Apply normalization once, then both attention and FFN in parallel
  • both: Normalize both before and after (not compatible with CrossAttention)

Documentation

Full documentation available at This Page

Contributing

Contributions are welcome!

License

Distributed under the Apache License 2.0. See LICENSE for more information.

Citation

If you use transformer in your research, please cite:

@software{transformer2026,
author = {Leinier Orama},
title = {transformer: PyTorch implementation of the current State-Of-The-Art(SOTA) Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/lof310/transformer}
}

About

PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages