Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neuron to Graph: Interpreting Language Model Neurons at Scale

Authors: Alex Foote, Neel Nanda, Esben Kran, Ionnis Konstas, Shay Cohen and Fazl Barez

Short paper accepted at RTML workshop at ICLR 2023: https://arxiv.org/abs/2304.12918

Longer preprint available here: https://arxiv.org/abs/2305.19911

Description

We provide a new interpretability tool for Language Models, Neuron to Graph (N2G). N2G builds an intepretable graph representation of any neuron in a Language Model, which can be visualised for human interpretation and used to simulate the behaviour of the target neuron by predicting activations on input text, allowing for a direct measure of the quality of the representation by comparing to the ground truth activations of the neuron. The resulting graphs are a searchable and programmatically comparable representation, facilitating greater automation of interpretability research.

Architecture
Overall architecture of N2G. Activations of the target neuron on the dataset examples are retrieved (neuron and activating tokens in red). Prompts are pruned and the importance of each token for neuron activation is measured (important tokens in blue). Pruned prompts are augmented by replacing important tokens with high-probability substitutes using DistilBERT. The augmented set of prompts are converted to a graph. The output graph is a real example which activates on the token "except" when preceded by any of the other tokens.

Given dataset examples that maximally activate a target neuron within a Language Model, N2G extracts the minimal sub-string required for activation, computes the saliency of each token for neuron activation, and creates additional examples by replacing important tokens with likely substitutes using DistilBERT. The set of enriched examples with token saliencies is then converted to a trie representing the tokens on which a neuron activates, as well as the context required for activation on these tokens. The trie can be used to process text and output token-level activations, which can be compared to the ground-truth activations of the neuron for automatic evaluation. A simplified version of this trie can also be visualised for human interpretation - activating tokens are coloured red according to how strongly they activate the neuron, and context tokens are coloured blue according to their importance for neuron activation. Once a model has been processed and a neuron graph has been generated for every neuron in the model, these graphs can be searched to identify neurons with particular properties, such as activating on a particular token when it occurs with another context token.

Examples

In-context
Neuron graph for an in-context learning neuron that activates on repeated token sequences. Identified by searching the graph representations for neurons which frequently have a repeated token in their context tokens as well as their activating tokens.
similar
A neuron graph that occurs for a neuron in Layer 1 and a neuron in Layer 4 of the model. The neurons have identical behaviour, and were recognised as a similar pair through an automated graph comparison process.
import
Neurons related to programming syntax, specifically import statements. Top - Neuron graph illustrating import syntax for the Go programming language. Bottom Left: Neuron graph showing fundamental elements of Python import syntax. Bottom Right: Neuron graph for a neuron that responds to the imports of widely-used Python packages.
superposition
A neuron graph exhibiting polysemanticity, with three disconnected subgraphs each responding to a phrase in a different language.

Citation

If you use N2G in your research, please cite one of our papers:

@inproceedings{foote2023neuron2graph, title={Neuron to Graph: Interpreting Language Model Neurons at Scale}, author={Foote, Alex and Nanda, Neel and Kran, Esben and Konstas, Ionnis and Cohen, Shay and Barez, Fazl}, booktitle={arXiv}, year={2023} } 

About

Tools for exploring Transformer neuron behaviour, including input pruning and diversification.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages