Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌐 TranslationLM — Neural Machine Translation (English → Italian)

LicensePythonPyTorchKaggle

A Transformer-based Neural Machine Translation model that translates English to Italian, trained from scratch on the opus_books corpus using the architecture described in "Attention Is All You Need" (Vaswani et al., 2017).


📋 Table of Contents


Overview

This project implements a complete neural machine translation pipeline — from raw text to trained model — without relying on pre-trained weights. The goal was to deeply understand the Transformer architecture by building and training every component from the ground up.

Key highlights:

  • Built the full Transformer encoder–decoder architecture in PyTorch
  • Trained BPE (Byte-Pair Encoding) tokenizers from scratch using HuggingFace tokenizers
  • Trained for 20 epochs on an NVIDIA Tesla P100 GPU (~4h 35m total)
  • Loss reduced from 6.17 → 2.28 (best checkpoint: 2.276 at epoch 18)
  • Qualitative translations improve visibly epoch-by-epoch (see Sample Translations)

Architecture

Input (English) Output (Italian)
│ ▲
▼ │
Token Embedding + Positional Encoding │
│ │
▼ ┌─────┴──────┐
┌──────────────┐ Encoder Output │ Projection │
│ Encoder │ ──────────────────►│ + Softmax │
│ (6 layers) │ └─────────────┘
└──────────────┘ ▲
│
┌──────┴──────┐
│ Decoder │
│ (6 layers) │
└─────────────┘
▲
│
Token Embedding + Positional Encoding
│
Target (Italian) shifted right

Each Encoder block consists of:

  1. Multi-Head Self-Attention (8 heads)
  2. Add & Norm (residual connection + layer normalisation)
  3. Position-wise Feed-Forward Network (d_ff = 2048)
  4. Add & Norm

Each Decoder block adds: 5. Masked Multi-Head Self-Attention (causal masking prevents looking ahead) 6. Multi-Head Cross-Attention (attends to encoder output)

Hyperparameters

ParameterValue
Model dimension (d_model)512
Feed-forward dimension (d_ff)2048
Attention heads (h)8
Encoder / Decoder layers (N)6
Dropout0.1
Max sequence length350
Batch size16
Learning rate1e-4 (Adam, ε=1e-9)
Label smoothing0.1
Epochs20
HardwareNVIDIA Tesla P100 (Kaggle)

Training Results

Training was conducted on the full opus_books en–it split (~29k sentence pairs, 90/10 train/val).

Loss Curve

EpochTrain LossEpochTrain Loss
16.173113.726
25.004123.738
35.172133.135
44.990143.423
54.682152.942
64.502163.075
74.372172.720
84.648182.276
93.900192.676
104.237202.287

Best checkpoint: Epoch 18 — loss = 2.276
Total training time: ~4 hours 35 minutes on Tesla P100
Steps per epoch: 1,819 @ ~2.22 batch/s

The non-monotonic loss curve (e.g., epochs 8, 10 slightly higher) is typical of Adam with label smoothing on small literary corpora — the model briefly overfits before generalising.


Project Structure

TranslationLM/
│
├── src/
│ ├── config.py # All hyperparameters and file path helpers
│ ├── model.py # Full Transformer architecture (from scratch)
│ ├── dataset.py # Bilingual dataset + BPE tokenizer training
│ ├── train.py # Training loop with TensorBoard + checkpointing
│ └── translate.py # Inference script (CLI)
│
├── notebooks/
│ └── translationlm.ipynb # Original Kaggle training notebook
│
├── requirements.txt
├── .gitignore
├── LICENSE # Apache 2.0
└── README.md

Quick Start

1. Clone & install

git clone https://github.com/atandra2000/TranslationLM.git
cd TranslationLM
pip install -r requirements.txt

2. Train the model

python src/train.py

This will:

  • Download the opus_books en–it corpus (~29k pairs) automatically
  • Train BPE tokenizers and save them as tokenizer_en.json / tokenizer_it.json
  • Save checkpoints to weights/translationlm_<epoch>.pt
  • Log loss curves to runs/ (view with tensorboard --logdir runs)

3. Translate a sentence

python src/translate.py --text "The sun sets over the mountains."# → Il sole tramonta sulle montagne.
python src/translate.py --text "She could not hide her feelings."# → Non riusciva a nascondere i suoi sentimenti.

4. View training curves (TensorBoard)

tensorboard --logdir runs
# Open http://localhost:6006

5. Run on Kaggle (original experiment)

See notebooks/translationlm.ipynb. The notebook mounts the source files as a Kaggle dataset and runs train.py directly on a P100 GPU.


Configuration

All hyperparameters live in src/config.py:

defget_config():
return {
'batch_size': 16,
'num_epochs': 20,
'learning_rate': 1e-4,
'seq_len': 350,
'd_model': 512,
'd_ff': 2048,
'h': 8, # attention heads'N': 6, # encoder/decoder layers'dropout': 0.1,
'datasource': 'opus_books',
'lang_src': 'en',
'lang_tgt': 'it',
}

Sample Translations

Greedy-decoded outputs logged during training on the same probe sentence show the model's progression from random noise to structured Italian:

EpochPredicted (greedy)
1Non si , ma non si , ma non si , e si , e si , e si .
4Per quanto a questo , egli , senza aver sentito , senza aver fatto il salotto ...
9Per quanto egli , senza pensare , si aspettava la fine ... dove non c'era nulla ...
13Per chiarire via , non si sentì , ma si aspettava la fine della sala di discorso ...
17Con questo sentimento , senza guardare , senza ascoltare , si aspettava la fine ...
20Con questo sentimento , senza guardare , uscì dalla fine di quella discussione e di quella sala , dove nessuno c'era ...

Source (probe):"To free himself from this feeling he went, without waiting to hear the end of the discussion, into the refreshment room, where there was no one except the waiters at the buffet."

Target:"Per liberarsi da questa sensazione penosa, senz'attendere la fine del dibattito, se ne andò in una sala dove non c'era nessuno, tranne i servitori vicino a una credenza."

The model learns Italian grammar structure by epoch 4–5, picks up key vocabulary by epoch 10, and produces largely fluent sentences by epoch 17–20.


Key Design Decisions

Why train from scratch?
The goal was to understand every component — not just call from_pretrained(). Building the attention mechanism, positional encoding, and training loop manually gave me a concrete understanding of what each piece does.

Why BPE tokenization?
Byte-Pair Encoding handles unknown words gracefully (by splitting into subword units) and produces compact vocabularies — important for a small corpus like opus_books.

Why label smoothing?
Label smoothing (ε=0.1) prevents the model from becoming overconfident on the training set, which helps generalisation on a small literary corpus.

Why greedy decoding during training?
Beam search would give better BLEU scores but greedy decoding is deterministic, fast, and sufficient for monitoring training quality per epoch.


Future Work

  • Implement beam search decoding for better translation quality
  • Evaluate BLEU score on a held-out test set
  • Add learning rate warm-up schedule (as in the original paper)
  • Extend to other language pairs (e.g., en–fr, en–de)
  • Experiment with larger datasets (WMT14, OPUS-100)
  • Deploy as a FastAPI web service

License

This project is released under the Apache 2.0 License.


Built and trained by Atandra Bharati

About

Transformer NMT from scratch for English→Italian — full encoder-decoder, word-level tokenizers (Whitespace + WordLevel) on opus_books, loss 6.17→2.28 over 20 epochs on P100

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages