Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); argon2: Add parallel lane processing by aumetra · Pull Request #149 · RustCrypto/password-hashes · GitHub
Skip to content

argon2: Add parallel lane processing - #149

Merged
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes
Apr 18, 2021
Merged

argon2: Add parallel lane processing#149
tarcieri merged 5 commits into
RustCrypto:masterfrom
aumetra:feature/parallel-lanes

Conversation

@aumetra

Copy link
Copy Markdown
Contributor

Adds parallel processing for lanes
Closes#103

The parallelism is gated behind the feature parallel

When the feature is activated, the #![forbid(unsafe_code)] gets downgraded to #![deny(unsafe_code)] due to unsafe usage (here)
The #![no_std] flag is being disabled as well

@tarcieri

Copy link
Copy Markdown
Member

Thanks for implementing this.

I think it might make sense to use rayon to manage the thread pool, similar to what we have in the pbkdf2 crate:

/// Generic implementation of PBKDF2 algorithm.
#[cfg(feature = "parallel")]
#[inline]
pubfnpbkdf2<F>(password:&[u8],salt:&[u8],rounds:u32,res:&mut[u8])
where
F:Mac + NewMac + Clone + Sync,
{
let n = F::OutputSize::to_usize();
let prf = F::new_varkey(password).expect("HMAC accepts all key sizes");
res.par_chunks_mut(n).enumerate().for_each(|(i, chunk)| {
pbkdf2_body(i asu32, chunk,&prf, salt, rounds);
});
}

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/lib.rs Outdated
Comment threadargon2/src/instance.rs
Comment threadargon2/src/lib.rs Outdated

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool, looks good to me.

I'll give @newpavlov a few days to comment in case he can think of a better solution to the mutable aliasing problem.

@tarcieri
tarcieri requested a review from newpavlovMarch 27, 2021 19:07

@nikomatsakisnikomatsakis left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I got intrigued by the twitter thread, as @Bascule hoped, but I don't quite understand what this code is trying to do. :( More pointers would be helpful! I do see some problems with it as written, though.

Comment threadargon2/src/instance.rs Outdated
Comment threadargon2/src/instance.rs
@tarcieri

tarcieri commented Mar 28, 2021

Copy link
Copy Markdown
Member

I can try to write a short synopsis of how the Argon2 paper describes parallel operation of the algorithm (mostly summarizing sections 3.2 and 3.3, see also section 6.2 Implementing parallelism):

https://www.password-hashing.net/argon2-specs.pdf

The algorithm operates over a matrix of "blocks" consisting of:

  • 𝒑 rows ("lanes"), where 𝒑 is the desired number of worker threads
  • 𝒒 columns

Each lane is further subdivided into 𝑺 = 4 "slices" (in Argon2 terminology, and referred to as SYNC_POINTS in the code), and the intersection of a slice and a lane forms a segment of length 𝒒 / 𝑺.

Segments of the same slice are computed in parallel and therefore cannot reference each other. So from a memory model perspective, what we'd really like is to mutably borrow the values of a particular slice, partition them into segments, and give each worker thread access to a particular segment.

However, we also need to allow all of the worker threads to simultaneously borrow all of the other blocks which do not belong to the "slice" being operated on to reference as inputs.

This is the tricky part: the fill_segment operation can reference blocks from the current lane, or other lanes, but will not reference blocks from the same "slice" being operated on.

Having just written all of that down (thanks for rubber ducking if nothing else), I think I have a better idea of how to model this problem safely in Rust: "slices" (in the Argon2 sense) should be the core level of granularity in which the working "memory" is organized.

The main loop of the algorithm iterates over the slices. Provided I'm actually understanding this correctly, we can borrow one slice mutably at the time and the others immutably. The mutably borrowed slice can then be subdivided into a segment for each lane, given to the worker threads along with immutable references to all of the other slices.

I think a big part of what's making this so tricky right now is the memory consists of a contiguous Vec<Block>. I think that might still be fine for the backing storage, but perhaps we could mediate access to segments through another type that splits the borrows by Argon2 "slice", allowing one "slice" to be acted on mutably and the others referenced as inputs.

@tarcieri

Copy link
Copy Markdown
Member

I think a next step which might be helpful in general is to extract a struct Memory which borrows from a backing [Block] slice, pass that to Instance::new instead of the raw &'a mut [Block], and provide a method on Memory for accessing blocks by their Position.

From there we can look at borrow splitting the backing buffer so Memory only holds 3 of the 4 "slices" at a time, and can be used to look up the "reference" blocks being used to fill a segment, but it would not have access to the slice being operated over (which would be borrowed mutably and split up into segments among the worker threads).

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also based on @nikomatsakis's comments I'm going to unapprove this for now

@aumetra
aumetra marked this pull request as draft March 29, 2021 13:38
@aumetra

Copy link
Copy Markdown
ContributorAuthor

This should fix the UB as every thread now dereferences the pointer itself

@tarcieri

Copy link
Copy Markdown
Member

@smallglitch nice! I think that's a start.

Do you want to mark the PR as ready for review?

@aumetra
aumetra marked this pull request as ready for review April 18, 2021 18:31
@aumetra

Copy link
Copy Markdown
ContributorAuthor

Sure, I wasn't sure if you'd be ok with an unsafe implementation
If an unsafe implementation is ok, I think this should be fine

@tarcieritarcieri left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using unsafe is fine for now. I can circle back on a safe implementation.

Thanks for extracting a Memory type.

@tarcieri
tarcieri merged commit dd8d13b into RustCrypto:masterApr 18, 2021
@tarcieri

Copy link
Copy Markdown
Member

I opened #154 to track making the implementation safe

@tarcieritarcieri mentioned this pull request Apr 18, 2021
@tarcieritarcieri mentioned this pull request Oct 2, 2021
dns2utf8 pushed a commit to dns2utf8/password-hashes that referenced this pull request Jan 24, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

argon2: parallel implementation

3 participants

@aumetra@tarcieri@nikomatsakis