Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Hamiltonian Deep Neural Networks

PyTorch implementation of Hamiltonian deep neural networks as presented in "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design".

Installation

git clone https://github.com/DecodEPFL/HamiltonianNet.git
cd HamiltonianNet
python setup.py install

Basic usage

2D classification examples:

./examples/run.py --dataset [DATASET] --model [MODEL]

where available values for DATASET are swiss_roll and double_moons.

Distributed training on 2D classification examples:

./examples/run_distributed.py --dataset [DATASET]

where available values for DATASET are swiss_roll and double_circles.

Classification over MNIST dataset:

./examples/run_MNIST.py --model [MODEL]

where available values for MODEL are MS1 and H1.

To reproduce the counterexample of Appendix III:

./examples/gradient_analysis/perturbation_analysis.py

Hamiltonian Deep Neural Networks (H-DNNs)

H-DNNs are obtained after the discretization of an ordinary differential equation (ODE) that represents a time-varying Hamiltonian system. The time varying dynamics of a Hamiltonian system is given by

and .

where y(t) ∈ ℝn represents the state, H(y,t): ℝn × ℝ → ℝ is the Hamiltonian function and the n × n matrix J, called interconnection matrix, satisfies .

After discretization, we have

  • H1-DNN:

  • H2-DNN:

where

2D classification examples

We consider two benchmark classification problems: "Swiss roll" and "Double circles", each of them with two categories and two features.

swissrolldoublecircles

An example of each dataset is shown in the figures above together with the predictions of a trained 64-layer H1-DNN (colored regions on the background). For these examples, the two features data is augmented, leading to yk ∈ ℝ4, k = 0,...,64.

Figures below shows the hidden feature vectors —the states yk— of all the test data after training. First, a change of basis is performed in order to have the classification hyperplane perpendicular to the first basis vector x1. Then, projections are performed on the new coordinate planes.

propagation Swiss rollpropagation Swiss rollpropagation Swiss roll

propagation Double circlespropagation Double circlespropagation Double circles

Counterexample

Previous work conjetured that some classes of H-DNNs avoid exploding gradients when y(t) varies arbitrarily slow. The following numerical example shows that, unfortunately, this is not the case.

We consider the simple case, where the underlying ODE is

(t) = ε J tanh( y(t) ) with .

We study the evolution of y(t) and yγ(t), t ∈ [t0, T] and t0 ∈ [0, T], with initial conditions y(t0) = y0 and yγ(t0) = y0 + γβ, with γ = 0.05 and β the unitary vectors. The initial condition y0 is set randomly, and normalized to have unitary norm.

y(t)_counterexamplephi(t)_counterexample

The left Figure shows the time evolution of y(t), in blue, and yγ(t), in orange, when a perturbation is applied at a time t0 = T-t. The nominal initial condition (y(T-t)) is indicated with a blue circle and the perturbated one (yγ(T-t)) with an orange cross. A zoom is presented on the right side, where a green vector indicates the difference between yγ(T) and y(T).

Figure on the right presents the entries (1,1) and (2,2) of the BSM matrix. Note that the value coincides in sign and magnitud with the green vector.

This numerical experiment confirms that the entries of the BSM matrix (we only show 2 of the 4 entries) diverge as the depth of the network increases (i.e. as the perturbation is introduced further away from the output).

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

The code DOI is DOI.

CC BY 4.0

References

[1] Clara L. Galimberti, Luca Furieri, Liang Xu and Giancarlo Ferrari Trecate. "Hamiltonian Deep Neural Networks Guaranteeing Non-vanishing Gradients by Design," arXiv:2105.13205, 2021.

[2] Clara L. Galimberti, Liang Xu and Giancarlo Ferrari Trecate. "A unified framework for Hamiltonian deep neural networks," The third annual Learning for Dynamics & Control (L4DC) conference, preprint arXiv:2104.13166 available, 2021.

[3] Eldad Haber and Lars Ruthotto. "Stable architectures for deep neural networks," Inverse Problems, vol. 34, p. 014004, Dec 2017.

[4] Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert and Elliot Holtham. "Reversible architectures for arbitrarily deep residual neural networks," AAAI Conference on Artificial Intelligence, 2018.

About

PyTorch implementation of Hamiltonian deep neural networks.

Resources

Stars

23 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages