Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

DeepDecipher

🦠 Official repository and open-source website for DeepDecipher.

Paper: Accessing and Investigating Neuron Activation in Large Language Models

Website available here. Contributers, see below for a setup guide. See the data available through the official API.

DeepDecipher is a package that exposes methods to generate information from arbitrary HookedTransformer classes, scrape existing databases of information about neurons within the field of mechanistic interpretability, search over neuron stores generated from the Neuron2Graph package, and set up a server with an API and a server with a UI that interfaces with the API.

As part of the publication of the paper, we also present a publicly available API developed using DeepDecipher (TBD).

See the search UI in action here: Search shows "he" returning 250+ results and "she" only returning about half

See the neuron information UI pages in action here: Showing a semantic graph of what the neuron activates to along with dataset examples that the neuron activates a lot to. Also shows that GPT-4 and similar neurons are not available.

Features

  • The DeepDecipher Python package to dynamically load neuron information from any available existing APIs, such as neuroscope.io and the OpenAI Neuron ExplainerAPI.
  • The DeepDecipher Python package also provides functionality to compile data folders from any setup script and serve it as a data-efficient API on a server. This includes a simple setup to use raw JSON.
  • The DeepDecipher API is an extensible and active API to relevant variables for single-neuron analysis
  • The API has access to relevant layer- and model-size information, such as layer neurons sorted by how interesting they are
  • The DeepDecipher front-end is an application to navigate the neurons in the style of neuroscope (Nanda, 2022)
  • We implement a search that reveal interesting examples of behavior

Data available per neuron

  • NeuroScope's max activating dataset examples on 25 models
  • Neuron2Graph's neuron activation model, along with the explanation power
  • GPT-4's neuron activation explanation, along with the explanation power

Future data ideas:

  • Which neurons have the most impact on this neuron's activation (based on weights)
  • The neuron's embedding based on Neuron2Graph model
  • Neuron interest variable: Variance / kurtosis of activation
  • Which neurons is it connected to within the MLP layers
  • Which neurons does this neuron impact the most (based on weights)
  • Which tokens it passes to the residual stream (?)
  • Neuron activation differences over training epochs (only available on Pythia models)
  • Most correlating neurons
  • Subnetwork analysis: Identification of groups of neurons that often activate together.
  • Topological role: Information about the neuron's role in the overall network topology (e.g., hub, peripheral, connector, etc.) using weighted directional network summary statistics methods
  • (?) Logit attribution: How much does this neuron affect the output

Data available per layer

  • Top interesting neurons
  • Links to all neurons
  • Meta data

Data available per model

  • Top interesting neurons by layer
  • Links to all layers
  • Meta data

JSON response

> print(request.get("https://apartresearch.com/DeepDecipher/api/GPT-2-XL/5/2332").json())
{
"model" : "GPT-2 XL",
"available" : ["Neuron Graph", "GPT-4 Explanation", "Max Activating Dataset Example"],
"layer" : 5,
"neuron" : 2332,
"metadata" : {
...
},
"neuroscope" : {
...
},
"neuron2graph" : {
"explanation-score" : 0.56,
...
},
"GPT-4" {
"explanation-score" : 0.43,
...
}
}

Contributor setup

This guide will ensure you have the right environment and start a small instance of DeepDecipher that serves only Neuroscope data on the solu-1l model. Tested in Windows Subsystem for Linux with Ubuntu 22.04.2 LTS.

  1. Ensure you have a working Python installation (at least version 3.7, tested with version 3.10.7).
  2. Ensure you have a working Rust toolchain (if you can use the cargo command it should be fine). See here to get one. Any version from the last few years should work. The newest one definitely will.
  3. Clone the repo and move to the root of the repo.
  4. Ensure a Python environment is active (conda, venv, whatever...)
  5. Install the maturin package by running python -m pip install maturin.
  6. Build the package by running maturin develop --release. The package will now be installed in your environment.
  7. Run the scrape Neuroscope script with python -m scripts.scrape_neuroscope. If this works, DeepDecipher is installed correctly. A file called data.db should be created in the root folder.
  8. Run python -m DeepDecipher data.db in the terminal to start the server.
  9. Visit http://localhost:8080/api/solu-1l/neuroscope/0/9 in the browser and you should see a JSON response with all the Neuroscope information on the 9th neuron of the solu-1l model.
  10. Navigate to http://localhost:8080/viz/solu-1l/all/0/9 and see various visualizations of the same neuron.

Screenshot of the frontend

Windows notes

On Windows, Maturin works less well, but there are workarounds.

  1. Make sure you clone the project into a path with no spaces.
  2. When building with Maturin, if you get the error Invalid python interpreter version or Unsupported Python interpreter, this is likely because Maturin fails to find your environment's interpreter. To fix this, instead of building with maturin develop, use maturin build --release -i py.exe (maybe replace py.exe with e.g. python3.exe if that is how you call Python) and then call python -m pip install .. The -i argument tells Maturin the name of the Python interpreter to use.

M1 notes

Problems arise when your Python version does not match your machine's architecture. This can happen on M1 chips since it is possible to run x86 Python even if the architecture is ARM. In this case, you can get an error that looks like

error[E0463]: can't find crate for `core`
|
= note: the `x86_64-apple-darwin` target may not be installed
= help: consider downloading the target with `rustup target add x86_64-apple-darwin

Simply download the x86 target with the suggested command and everything should work.

Models available

ModelInitialisationActivation FunctionDatasetLayersNeurons per LayerTotal NeuronsParameters
solu-1lRandomsolu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
gelu-1lRandomgelu80% C4 (Web Text) and 20% Python Code12,0482,0483,145,728
solu-2lRandomsolu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
gelu-2lRandomgelu80% C4 (Web Text) and 20% Python Code22,0484,0966,291,456
solu-3lRandomsolu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
gelu-3lRandomgelu80% C4 (Web Text) and 20% Python Code32,0486,1449,437,184
solu-4lRandomsolu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
gelu-4lRandomgelu80% C4 (Web Text) and 20% Python Code42,0488,19212,582,912
solu-6lRandomsolu80% C4 (Web Text) and 20% Python Code63,07218,43242,467,328
solu-8lRandomsolu80% C4 (Web Text) and 20% Python Code84,09632,768100,663,296
solu-10lRandomsolu80% C4 (Web Text) and 20% Python Code105,12051,200196,608,000
solu-12lRandomsolu80% C4 (Web Text) and 20% Python Code126,14473,728339,738,624
gpt2-smallRandomgeluOpen Web Text123,07236,86484,934,656
gpt2-mediumRandomgeluOpen Web Text244,09698,304301,989,888
gpt2-largeRandomgeluOpen Web Text365,120184,320707,788,800
gpt2-xlRandomgeluOpen Web Text486,400307,2001,474,560,000
solu-1l-pileRandomsoluThe Pile14,0964,09612,582,912
solu-4l-pileRandomsoluThe Pile42,0488,19212,582,912
solu-2l-pileRandomsoluThe Pile22,9445,88812,812,288
solu-6l-pileRandomsoluThe Pile63,07218,43242,467,328
solu-8l-pileRandomsoluThe Pile84,09632,768100,663,296
solu-10l-pileRandomsoluThe Pile105,12051,200196,608,000
pythia-70mRandomgeluThe Pile62,04812,28818,874,368
pythia-160mRandomgeluThe Pile123,07236,86484,934,656
pythia-350mRandomgeluThe Pile244,09698,304301,989,888

Repo standards

We use the Gitmoji commit standards.

References

To cite our work, please use the following BibTeX entry:

@misc{garde2023deepdecipher,
title={DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models},
author={Albert Garde and Esben Kran and Fazl Barez},
year={2023},
eprint={2310.01870},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Releases

Packages

Used by

Contributors

Languages