Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Official Repository for Understanding Integer Addition in Transformers, ICLR 2024 | Authors Philip Quirke and Fazl Barez

Understanding the inner workings of machine learning models like Transformers is vital for their safe and ethical use. This repository contains a CoLab ( ./Understanding_Addition_in_Transformers.ipynb ) that presents an in-depth analysis of a one-layer Transformer model trained for integer addition.

The CoLab creates, trains and then analyses a 1 layer Transformer model that performs integer addition e.g. 33357+82243=115600. Each digit is a separate token. For 5 digit addition, the model is given 12 "question" (input) tokens, and the model must then predict the correct 6 "answer" (output) tokens:

QuestionAnswer

This CoLab allowed the authors to understand the model's addition algorithm, summarised in this diagram:

StaircaseA3_Summary

Summary diagram legend: A: The 5 digit question is revealed token by token. The highest-value digit is revealed first. B: From the "=" token, the model attentions heads focus on successive pairs of digits, giving a 'staircase' attention pattern visible in 15 digit, 5 digit, etc addition. C: The 3 heads are time-offset from each other by 1 token so in each epoch data from 3 tokens is available. D: To calculate A3, the 3 heads do independent simple mathematical calculations on D3, D2 & D1. The results are combined by the MLP layer using 60 trigrams. A3 is calculated one token before it is needed. This approach is repeated for all answer digits.

This more detailed diagram shows how the model's addition algorithm is "coded" in the attention heads: StaircaseA3_Detailed

Detailed diagram legend: A: For A3, the addition algorithm combines information from digits 3, 2 and 1. B: 1st Head calculates MC1 on digit 1. C: 2nd Head calculates MC1 and MS9 (which are independent of each other and so at most one is true) on digit 2. D: 3rd Head does Base Add on digit 3. E: The MLP layer uses trigrams to combine the information from the 3 heads to give the final answer A3.

Tips for using the Colab

  • You can run all the code in the CoLab notebook yourself in Google CoLab ( https://colab.research.google.com/ ). You can change the code and experiment.
  • To use the notebook, in Google CoLab, you will need to go to Runtime > Change Runtime Type and select GPU as the hardware accelerator.
  • The graphs are interactive!
  • Use the table of contents pane in the sidebar to navigate
  • Collapse irrelevant sections with the dropdown arrows
  • Search the page using the search in the sidebar, not CTRL+F

Reference

If you use or find this work helpful, please consider citing our work:

@inproceedings{quirke2024understanding,
title={Understanding Addition in Transformers},
author={Philip Quirke and Fazl Barez},
booktitle={International Conference on Learning Representations (ICLR)},
year={2024},
address={Vienna, Austria},
eprint={2310.13121},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

About

✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks

Resources

Stars

19 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages