Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - lhez/qualcomm · GitHub
Skip to content

Repository files navigation

Qualcomm Adreno TVM Evaluation Repo

***Disclaimer: This is a development repository, texture memory support is currently in the upstreaming process and has been cut from the TVM subtree contained herein. See the TVM discuss forum RFC for more information and links to relevant PRs.

Last version of TVM this was evaluated on and worked (01/28/2021): 4abbe4902e451cc5a963b8b60a70e548d48ace62.

For testing texture memory support, please use the tvm repository included as a subtree in this repository: tvm.

Questions of issues using the scripts? Submit a ticket via the OctoML helpdesk.

Testing model performance with texture memory:

In the table below you can see the performance numbers (inference time in milliseconds) which were achieved on the Realme GT 5G.

mace_mobilenetv1mace_resnet50_v2mace_inceptionv3vgg16mace_deeplabv3mace_yolov3
TVM textures FP164,827,4644,4578,96103,99175,09
TVM textures FP16a325,4228,5144,3196,64110,32242,34
TVM textures FP327,6640,5669,87131,99154,06306,27

The tuning log files are located in logs/mace_models/. You can use the evaluate.py script for reproducing these numbers. Copy the name of the model from the table and use the relevant log file with tuned statistic. Below, you can see examples of run mobilenetv1:

# float16 compute, float16 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc16.autotvm.log
# float16 compute, float32 accumulate
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float16 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float16.acc32.autotvm.log
# float32 inference
python ./scripts/evaluate.py -m mace_mobilenetv1 -t float32 -k android --target="opencl --device=adreno" -l ./logs/mace_models/mace_mobilenetv1.texture.float32.acc32.autotvm.log

Refer to the below instructions for running the scripts/evaluate.py script for more information

Running texture.py tests:

scripts/texture.py is a set of compute and schedule definitions for various workloads employing texture memory cache stage when the -m "texture" argument is supplied. For each test, numerical comparisons are checked against numpy results. Some of the tests can be tuned with the --tune flag. Log files with autotvm tuning records exist in the logs/ directory for many these tunable tests. See the below for a few invocation examples on how to run a tuned schedule with texture memory.

usage: scripts/texture.py [-h] [-m MEMORY] [-s] [-l LOG] [-T] -t TEST
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-k RPC_KEY]
Set test arguments
optional arguments:
-h, --help show this help message and exit
-m MEMORY, --memory MEMORY
Use global or texture
-s, --shared Use shared memory
-l LOG, --log LOG AutoTVM tuning record logfile
-T, --tune Whether to tune or not
-t TEST, --test TEST Selected test to run
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use

Example invocations,

# ------------------------
# Conv2d VGG16 layer [3x3]
# ------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.autotvm.shared.log
> 115.4 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.shared.autotvm.best.log -m texture -s
> 116.9 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune2 -m texture -l logs/conv2d_NCHWc_KCRSk_tx_tune2.texture.noshared.autotvm.log
> 147.6 GFLOPS
# ------------------------------
# Conv2d MobilenetV1 layer [1x1]
# ------------------------------
# Memory hierarchy: shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s
> 100.2 GFLOPS
# Memory hierarchy: texture->shared->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune_1024.log -s -m "texture"
> 89.2 GFLOPS
# Memory hierarchy: texture->local
$ python scripts/texture.py -r 0.0.0.0 -p 9191 -k android --test=conv2d_NCHWc_KCRSk_tx_tune -l logs/conv2d_NCHWc_KCRSk_tx_tune.texture.noshared.log -m "texture"
> 137.5 GFLOPS

Setting up the host development machine

On the host machine (typically your development box) you'll need to build TVM.

git clone https://github.com/octoml/qualcomm --recursive
cd qualcomm/tvm
mkdir build
cp cmake/config.cmake build/.
echo 'set(USE_LLVM llvm-config)' >> build/config.cmake
echo 'set(USE_GRAPH_RUNTIME_DEBUG ON)' >> build/config.cmake
cd build
cmake ..
make -j8
cd ..
export TVM_HOME=$PWD
export PYTHONPATH=$TVM_HOME/python:${PYTHONPATH}
export LD_LIBRARY_PATH=${TVM_HOME}/build:$LD_LIBRARY_PATH

Cross compiling the C++ RPC server for Android

Refer to the documentation here to cross compile the C++ RPC binary and tvm_runtime libraries for Android.

To run, use adb to push the cross compiled tvm_rpc binary and libtvm_runtime.so shared library to /data/local/tmp on the Android device. Then run the RPC server with:

adb shell
cd /data/local/tmp
LD_LIBRARY_PATH=. ./tvm_rpc server --tracker=<tracker IP>:<tracker port> --key=android

Setting up the RPC device tracker

Once TVM is built on the host, you'll need to launch the RPC tracker service with the following command:

python -m tvm.exec.rpc_tracker --host=<tracker IP> --port=<tracker port> --port-end=9192

Where tracker IP is the host IP, and tracker port can be 9191.

When done, you can register the Android device on the tracker with the same key used to run the on device RPC server:

python -m tvm.exec.rpc_server --tracker <tracker host>:<tracker port> --key android

Finally, make sure that the hardware is properly registered to the tracker. On the host, or any machine connected to the local network, check the devices registered on the tracker with the following command:

python -m tvm.exec.query_rpc_tracker --host <tracker IP> --port <tracker port>

Using the experiment script

Under scripts you'll find a python script evaluate.py that can evaluate or tune a set of models:

Below is the usage for the script, which you can get with

$ python3 scripts/evaluate.py -h
usage: evaluate.py [-h] -m
{resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
[-t {float32,float16}] [-l LOG] [-k RPC_KEY]
[-r RPC_TRACKER_HOST] [-p RPC_TRACKER_PORT] [-T TARGET]
[--tune TUNE] [--debug DEBUG]
Tune and/or evaluate a curated set of models
optional arguments:
-h, --help show this help message and exit
-m {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}, --model {resnet50,mobilenetv1,inceptionv3,vgg16,mobilenetv3_ssdlite,deeplabv3}
Model to tune and/or evaluate
-t {float32,float16}, --type {float32,float16}
Specify whether the model should be run with single or
half precision floating point values
-l LOG, --log LOG AutoTVM tuning logfile name
-k RPC_KEY, --rpc_key RPC_KEY
RPC key to use
-r RPC_TRACKER_HOST, --rpc_tracker_host RPC_TRACKER_HOST
RPC tracker host IP address
-p RPC_TRACKER_PORT, --rpc_tracker_port RPC_TRACKER_PORT
RPC tracker host port
-T TARGET, --target TARGET
Compilation target
--tune TUNE Whether or not to run autotuning
--debug DEBUG Use graph runtime debugger to output per layer perf.
data and other statistics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages