[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[XNNPACK][Weights Cache] Enable in XNNPACK - #9155

Merged
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head
Mar 14, 2025
Merged

[XNNPACK][Weights Cache] Enable in XNNPACK#9155
facebook-github-bot merged 9 commits into
gh/mcr229/11/basefrom
gh/mcr229/11/head

Conversation

@mcr229

@mcr229mcr229 commented Mar 11, 2025

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

We enable the XNNPACK Weights cache in XNNPACK.

the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).

Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.

In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.

After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.

Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.

We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_

Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Mar 11, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9155

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit c4c62de with merge base 630d0cc (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mcr229 added a commit that referenced this pull request Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
ghstack-source-id: 271070693
Pull Request resolved: #9155
@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 11, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271090604
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@mcr229mcr229 added the release notes: xnnpack Changes to the XNNPack backend delegate label Mar 11, 2025
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 11, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271095503
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 12, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271379632
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 13, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271658387
@exported-using-ghexport
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271732050
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271809922
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271819431
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
[ghstack-poisoned]
mcr229 added a commit that referenced this pull request Mar 14, 2025
Pull Request resolved: #9155
We enable the XNNPACK Weights cache in XNNPACK.
the weights cache is initialized for the runtime with the named data map and a memory allocator (for now the memory allocator is not used, but i hope in the future this can be used to managed the memory for packed weights).
Before Creating the runtime, we first initialize the weights cache, this sets the finalization state to false. As we add weight/bias tensors to the graph, we load them through the named data map in the weights cache, and keep a map of the pointer to the name. When XNNPACK Creates the runtime and packs the weights, it uses the weights_cache method look_up_or_insert. We use the pointers provided in the cache key to look up their names and append them together like ("weightsbias"). We then insert the packed weights with that key.
In future look ups, we just use the pointer cached at the named pack tensor key, saving us from packing in the future.
After creating the runtime and packing the weights, we finalize the cache. This sets is_finalized to true. We also free all unpacked buffers loaded from the named data map as they are no longer needed. We also keep reference counts for all the packed weights incrementing the packed weights which were used by this runtime. We return a vector of all the packed weight names to the xnn_executor runner. When the XNNExecutor is destroyed, we decrement the counts of the packed buffers and destroy them if necessary.
Note that this feature is gated behind the XNN_ENABLE_WEIGHTS_CACHE flag. Since the weights_cache is a global member of the singleton xnnpack backend class, and it is also read/write, we add a mutex to ensure that access to the weights_cache is thread safe.
We added a new mutex, so the mutex hiearchy is:
workspace_mutex_ -> weights_cache_mutex_
ghstack-source-id: 271823386
@exported-using-ghexport
Internal:
I ran a simple experiment with the Machine translation model. I loaded encode_first executed it, and then loaded forward and executed (both methods staying in memory). I measured the RSS after Encode First and then measured the RSS after forward. We saw the following results:
| | RSS after Encode First (MiB) | RSS after Forward (MiB) |
|----------------------|------------------------------|-------------------------|
| Without Weight Cache | 62.765625 | 130.019531 |
| With Weight Cache | 62.789062 | 93.222656 |
Which shows that with the weight cache and two methods, we can see around ~28% reduction in memory usage
Differential Revision: [D70885926](https://our.internmc.facebook.com/intern/diff/D70885926/)
@facebook-github-bot

Copy link
Copy Markdown
Contributor

This pull request was exported from Phabricator. Differential Revision: D70885926

@facebook-github-bot
facebook-github-bot merged commit cbd4a0f into gh/mcr229/11/baseMar 14, 2025
@facebook-github-bot
facebook-github-bot deleted the gh/mcr229/11/head branch March 14, 2025 21:26
SS-JIA pushed a commit that referenced this pull request Mar 15, 2025
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #9155 by
@mcr229
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/mcr229/10/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/mcr229/11/orig
@diff-train-skip-merge
---------
Co-authored-by: Max Ren <maxren@meta.com>
@SS-JIA
SS-JIA restored the gh/mcr229/11/head branch March 15, 2025 02:54
@SS-JIA
SS-JIA deleted the gh/mcr229/11/head branch April 16, 2025 20:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.fb-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@mcr229@facebook-github-bot@kirklandsign