Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

FastSparse

Customizable Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL).

Warning: this repo is undergoing active development

Getting Started

Install

pip install fastsparse

Sparse Algorithms

This network implements the following sparse algorithms:

Abbr.Sparse Algorithmin FastSparseNotes
static sparsity baselineomit DynamicSparseTrainingCallback
SETSparse Evolutionary Training (Jan 2019)DynamicSparseTrainingCallback(**SET_presets)
SNFSSparse Networks From Scratch (Jul 2019)DynamicSparseTrainingCallback(**SNFS_presets)*redistribution not implemented
RigLRigged Lottery (Nov 2019)DynamicSparseTrainingCallback(**RigL_presets)

*Authors of the RigL paper demonstrate that using SNFS + Erdos-Renyi-Kernel distribution - redistribution outperforms SNFS + uniform sparsity + redistribution (at least on the measured benchmarks).

Fastai demo

With just 4 additional lines of code, you can train your model using the latest dynamic sparse training techniques. This example achieves >99% accuracy on MNIST using a ResNet34 with only 1% of the weights.

# (0) install the library# ! pip install fastsparse fromfastai.vision.allimport*# (1) import this packageimportfastsparseassparsepath=untar_data(URLs.MNIST)
dls=ImageDataLoaders.from_folder(path, 'training', 'testing')
learn=cnn_learner(dls, resnet34, metrics=error_rate, pretrained=False)
# (2) sparsify initial model + enforce maskssparse_hooks=sparse.sparsify_model(learn.model, model_sparsity=0.99,
sparse_f=sparse.erdos_renyi_sparsity)
# (3) schedule dynamic mask updatescbs= [sparse.DynamicSparseTrainingCallback(**sparse.SNFS_presets, batches_per_update=32)]
learn.fit_one_cycle(5, cbs=cbs)
# (4) remove hooks that enforce maskssparse_hooks.remove()

Simply omit the DynamicSparseTrainingCallback to train a fixed-sparsity model as a baseline.

PyTorch demo (not implemented yet)

importtorchfromtorchvisionimportmodelsdata= ...
model= ...
opt= ...
opt=DynamicSparseTrainingOptimizerWrapper(model, opt, **RigL_kwargs)
### Modified training step# sparse_opt.step(...) will determine whether to:# (A) take a regular opt step, or# (B) update network connectivitydefsparse_train_step(model, xb, yb, loss_func, sparse_opt, step, pct_train):
preds=model(xb)
loss=loss_func(preds, yb)
loss.backward()
sparse_opt.step(step, pct_train)
sparse_opt.zero_grad()

Save/Reload demo

Here is an example of saving a model and reloading it to resume training.

fromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity)
dst_kwargs= {**SNFS_presets, **{'batches_per_update': 8}}
cbs=DynamicSparseTrainingCallback(**dst_kwargs)
learn.fit_flat_cos(5, cbs=cbs)
# (0) save model as usual (masks are stored automatically)save_model('sparse_tiny_mnist', learn.model, learn.opt)
epochtrain_lossvalid_lossaccuracytime
00.3341190.6796190.50500700:03
10.2716450.5551700.84835500:02
20.2371150.0720880.97854100:02
30.2205530.0449270.98712400:02
40.1745850.0064961.00000000:02
### manually restart notebook #### (1) then recreate learner as usualfromfastai.vision.allimport*fromfastsparseimport*path=untar_data(URLs.MNIST_TINY)
dls=ImageDataLoaders.from_folder(untar_data(URLs.MNIST_TINY))
learn=cnn_learner(dls, resnet18, metrics=accuracy, pretrained=False)
# (2) re-sparsify model (this adds the masks to the parameters)sparse_hooks=sparsify_model(learn.model, model_sparsity=0.9, sparse_f=erdos_renyi_sparsity) # <-- initial sparsity + enforce masks# (3) load model as usualload_model('sparse_tiny_mnist', learn.model, learn.opt)
# (5) check validation loss & accuracy to verify we've loaded it successfullyval_loss, val_acc=learn.validate()
print(f'validation loss: {val_loss}, validation accuracy: {val_acc}')
# (4) optionally, continue training; otherwise remove sparsity-preserving hookssparse_hooks.remove()
/home/dc/anaconda3/envs/fastai/lib/python3.8/site-packages/fastai/learner.py:53: UserWarning: Could not load the optimizer state.
if with_opt: warn("Could not load the optimizer state.")
validation loss: 0.006496043410152197, validation accuracy: 1.0

Training with Large Batch Sizes

Authors of the Rigged Lottery paper hypothesize that the effectiveness of using the gradient magnitude for determining which connections to grow is partly due to their large batch size (4096 for ImageNet). Those without access to multi-gpu clusters can achieve effective batch sizes of this size by using fastai's GradientAccumulation callback, which has been tested to be compatible with this package's DynamicSparseTrainingCallback.

Training with Small # of Epochs

Dynamic sparse training algorithms work by modifying the network connectivity during training, dropping some weights and allowing others to regrow. By default, network connectivity is modified at the end of each epoch. When training with few epochs, however, there will be few chances to explore which weights to connect. To update more frequently, in DynamicSparseTrainingCallback, set batches_per_update to a smaller # of batches than occur in one training epoch. Varying the number of batches per update trades off the frequency of updates with stability in making good updates.

Customization

There are many ways to implement and test your own dynamic sparse algorithms using FastSparse.

Custom Initial Sparsity Distribution:

Define your own initial sparsity distribution by setting sparsify_method in sparsify_model to a custom function. For example, this function (included in library) will keep the first layer dense and set the remaining layers to a fixed sparsity.

deffirst_layer_dense_uniform(params:list, model_sparsity:float):
sparsities= [1.] + [model_sparsity] * (len(params) -1)
returnsparsities

Custom Drop Criterion

While published papers like SNFS and RigL refer to 'drop criterion', this library implements the reverse, a 'keep criterion'. This is a function that returns a score for each weight, where the largest M scores will be and M is determined by the decay schedule. For example, both Sparse Networks From Scratch and Rigged Lottery both use the magnitude of the weights (in FastSparse: weight_magnitude).

This can easily be customized in FastSparse by defining your own keep score function:

defcustom_keep_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., keep_score_f=custom_keep_scoring_function)

Custom Grow Criterion

The grow criterion is a function that returns a score for each weight, where the largest N scores will be and N is determined by the decay schedule. For example, Sparse Networks From Scrath grows weights according to the momentum of the gradient, while Rigged Lottery uses the magnitude of the gradient (in FastSparse, gradient_momentum and gradient_magnitude respectively).

defcustom_grow_scoring_function(param, opt):
score= ...
assertparam.shape==score.shapereturnscore

Then pass your custom function into the sparse training callback:

DynamicSparseTrainingCallback(..., grow_score_f=custom_grow_scoring_function)

Replication Results

In machine learning, is very easy for seemingly insignificant differences in algorithmic implementation to have a noticeable impact on final results. Therefore, this section compares results from this implementation to results reported in published papers.

TODO...

Under-The-Hood Details

Here's what's going on.

When you run sparsify_model(learn.model, 0.9), this adds sparse masks and add pre_forward hooks to enforce masks on weights during forward pass.

By default, a uniform sparsity distribution is used. Change the sparsity distribution to Erdos-Renyi with sparsify_model(learn.model, 0.9, sparse_init_f=erdos_renyi), or pass in your custom function (see Customization

To avoid adding pre_forward hooks, use sparsify_model(learn.model, 0.9, enforce_masks=False).

When you add the DynamicSparseTrainingCallback callback, ... TODO complete section

About

Fastai+PyTorch implementation of sparse model training methods (SET, SNFS, RigL) + customize-your-own.

Topics

Resources

Contributing

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

Generated from fastai/nbdev_template