Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

tensor_parallel

PyPI versionBlackCI status

🚀 Try new 40B LLMs demo in Kaggle

Run large PyTorch models on multiple GPUs in one line of code with potentially linear speedup.

importtransformersimporttensor_parallelastptokenizer=transformers.AutoTokenizer.from_pretrained("facebook/opt-13b")
model=transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b") # use opt-125m for testingmodel=tp.tensor_parallel(model, ["cuda:0", "cuda:1"]) # <- each GPU has half the weightsinputs=tokenizer("A cat sat", return_tensors="pt")["input_ids"].to("cuda:0")
outputs=model.generate(inputs, num_beams=5)
print(tokenizer.decode(outputs[0])) # A cat sat on my lap for a few minutes ...model(input_ids=inputs, labels=inputs).loss.backward() # training works as usual

Installation

Latest stable version (recommended):

pip install tensor_parallel

Bleeding edge version:

pip install https://github.com/BlackSamorez/tensor_parallel/archive/main.zip

Usage

Simply wrap your PyTorch model with tp.tensor_parallel and use it normally. For best memory efficiency, call tp.tensor_parallel while the model is still on CPU.

Here are a few use cases:

Advanced parameters to tensor_parallel:

  • device_ids: List[device] - which devices to use; defaults to all available GPUs
  • output_device: device - model outputs will have this device
  • tensor_parallel_config: tp.Config - use custom parallelism strategy, see slicing_configs.py
  • distributed: bool - if True, use torch.distributed backend instead of threading (requires torchrun)
  • sharded: bool - if True, find all trainable parameters that weren't split by Tensor Parallelism and split them using ZeRO-3 algorithm.
    • weights will be split between GPUs and re-assembled before each forward pass
    • TL;DR use this when training to avoid duplicate parameters (enabled by default!)
    • sharded_param_names: List[str] - parameter names that should be sharded this way, default = found automatically

Saving the model

To save a model such that it could be used in a non tensor_parallel context, you should use a save_tensor_parallel context wrapper.

importtorchimporttransformersimporttensor_parallelastpmodel=tp.tensor_parallel(
transformers.AutoModelForCausalLM.from_pretrained("facebook/opt-13b"), )
# A whole lot of trainig...withtp.save_tensor_parallel(model):
torch.save(model.state_dict(), "/tmp/")
# or model.save_pretrained("/tmp/")

Such code saves a model as if it was never split. It works by gathering model parts during state_dict creation.

Memory efficient dispatch

Normally, to normally create and dispatch a tensor_parallel model, one needs the whole model in memory. This can be troublesome, but there is another way.

It's possible to convert a state_dict of a basic model into the corresponding tensor_parallelstate_dict using a helper function convert_state_dict. The state dict can then be dispatched and loaded into the model:

importaccelerateimporttransformersimporttensor_parallelastp# Initialize a weightless tensor_parallel model from MyModelwithaccelerate.init_empty_weights():
model=tp.TensorParallel(
MyModel(),
device_ids=[0, 1] # and prepare it to be put on GPUs 0 and 1
)
# Load partial state_dict for MyModelstate_dict=torch.load("my_model_part_1_of_5.bin")
# Convert it into a tensor_parallel state_dicttensor_parallel_state_dict=tp.convert_state_dict(
state_dict,
tensor_parallel_config=model.tensor_parallel_config,
world_size=len(model.devices),
)
# Dispatch the partial state_dict (load_state_dict doesn't work with meta so here I use accelerate)device_map=tp.infer_sharded_device_map(model)
forparam_name, paraminstate_dict.items():
module_name=param_namewhilelen(module_name) >0andmodule_namenotindevice_map:
module_name=".".join(module_name.split(".")[:-1])
param_device=device_map[module_name]
accelerate.utils.set_module_tensor_to_device(model, param_name, param_device, value=param)

With this no more than one part of the model needs to be loaded into memory at once.

FAQ

  • Q: I don't have a multi-GPU server. Can I use tensor_parallel in Google Colab?

  • A: Colab has a single GPU, so there's no point in tensor parallelism. However, Kaggle offers two T4 for free to all phone-verified accounts.

  • Q: What is tensor parallelism?

  • A: You split each layer's weights into parts, multiply each part on a separate GPU, then gather results. Read more here

  • Q: Should I use TensorParallel or DataParallel?

  • A: TensorParallel for large models, DataParallel for smaller ones

  • Q: How does it compare against FullyShardedDataParallel and ZeRO?

  • A: ZeRO is better if you can fit a large batch, TensorParallel is better for small batches

Why use tensor_parallel ...

  • v.s. DeepSpeed and FairScale
    • DeepSpeed has many parallelization strategies, but requires careful configuration
    • tensor_parallel has one strategy that works with 1 line of code
    • tensor_parallel works in a jupyter notebook
  • v.s. MegatronLM
    • MegatronLM has great tensor parallelism for one model architecture
    • tensor_parallel has good parallelism for any architecture
    • tensor_parallel is way easier to install
  • v.s. parallelformers
    • parallelformers is inference-only, tensor_parallel supports training
  • v.s. alpa
    • alpa is a powerful tool for automatic distributed training / inference in JAX
    • tensor_parallel works with PyTorch
  • v.s. Model.parallelize()
    • both are easy to use, both fit large models
    • in parallelize, one GPU works at a time
    • in tensor_parallel, GPUs work in parallel

In short, use tensor_parallel for quick prototyping on a single machine. Use DeepSpeed+Megatron or alpa for million-dollar training runs.

Troubleshooting

If you experience NCCL errors, or random hanging, you may have some code errors that are not displayed properly. To debug these errors, we recommend restarting with export TENSOR_PARALLEL_USE_NATIVE=1 or on a single device.

If you found a bug or encountered a problem, please report it to our issue tracker. We will do our best to help, but it may take some time before we get to it. Please create issues only if your problem is specifically with tensor_parallel. For example, if you need help installing transformers or optimizing your code, please seek it elsewhere.

Code style

We use black and isort for all pull requests. Before committing your code, simply run black . && isort . and you will be fine.


About

Automatically split your PyTorch models on multiple GPUs for training & inference

Topics

Resources

Stars

654 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages