Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Installation

pip install -r requirements.txt

Then, in the vSim directory, run

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

Docker image

You may try our artifact using the docker image

# create image
docker build -t vsim:0.1 .# start container
docker run --name ndss26ae -it -d vsim:0.1
# use the bash
docker exec -it ndss26ae bash
cd /vSim &&source source.sh

Preprocess with IDA

The script ida/ida_script.py was tested with IDA v7.5.

<ida_path> -A -S"ida/ida_script.py <output_path>" bin_path

The output_path stores the extracted binary function entries.

For stripped binaries, vSim depends on IDA to identify the function entries.

Value Extraction

To run our customized symbolic execution engine, please run the following command. Please remember the dump_dir and cache_dir. In dump_dir, you will get many <func_addr>.pkl files. In cache_dir, there is a *.pkl file, which stores the preprocessed binary information.

# use ida's binary function entries
python ./src/bin_analyzer.py \
--bin <binary> \
--dump <dump_dir> \
--cache <cache_dir> \
--ida <ida_pkl_path> \
--loglevel 0 \
--workers <workers>

Fingerprint Generation

This process combines our Value Filtering, Normalization, Concretization.

Please prepare a CSV file, with columns binary,dump_dir,cache_dir,bin_id. binary, dump_dir, and cache_dir are the absolute paths used while value extraction. bin_id is the identifier of this binary. The binaries compiled from the same source should have the identical bin_id.

Please modify the paths for csv_path and fingerprint_dir.

  • csv_path is the aforementioned CSV file.
  • fingerprint_dir is the directory to store all generated fingerprints
fromutilsimport*importloggingfromsrc.pool.poolimportBinaryPoolfromsrc.bin_analyzerimportBinAnalyzerfromsrc.sim.baseimportSimBasecsv_path='xxx'fingerprint_dir='yyy'func_pool_path=csv_path+'.pkl'expr_timeout=30workers=32e_workers=1mode='refined'if__name__=='__main__':
expr_logger=get_logger_for_file(logging.ERROR, 'expr_sim')
get_logger_for_file(logging.ERROR, logname='binary_pool')
binary_pool=BinaryPool(csv_path, fingerprint_dir)
ifos.path.exists(func_pool_path):
func_pool=load_pkl(func_pool_path)
else:
func_pool=binary_pool.build_function_pool(
workers=workers,
expr_vec_workers=e_workers,
mode=mode,
expr_timeout=expr_timeout)
dump_pkl(func_pool, csv_path+'.pkl')

Function pool construction and similarity analysis

The above step have built pool of binaries. Given two bianry pools (e.g., gcc-O0 pool and gcc-O3 pool), we can perform pairwise match between them. Run the following code will perform Fingerprint Propagation, build pandas Dataframe as function pools, and compute their similarity pairwisely.

fromscripts.pool_pairwise_matchimportSETTINGS, get_ppw, runcsv1='xxx'csv2='yyy'fingerprints_dir='zzz'pool_size=10000# it should be false if you compare these two binary pools first time# Then, you can set it to True and always compare the same function pairsload_existing_key=Falsesetting='all_features'top_k, mrr, cost=run(csv1, csv2,
fingerprints_dir, pool_size, setting,
load_existing_key=load_existing_key,
key_path=key_file_path,
with_weights=True, workers=64)
res.append({'db_csv': csv1,
'input_csv': csv2,
'setting': setting,
'top_k': top_k,
'K_list': [1, 3, 5, 10, 50, 100],
'mrr': mrr,
'MRR_K_list': [None, 10, 100],
'pool_size': pool_size,
'cost': cost})
print(csv1, csv2, top_k, mrr)

Guidance for reproduce results

Assume the unzip vSim directory has a path $VSIM_HOME.

First of all, ensure two environment variables have been set

export PYTHONPATH=$(pwd)export VSIM_HOME=$(pwd)

In the docker image, running cd /vSim && source source.sh has the same effect.

Preparation

Please download the zip files for datasets and test first.

Please enter $VSIM_HOME/test folder and unzip all zip files:

unzip cross-optimization.zip
unzip cross-compiler.zip
unzip GMN-dataset-2.zip
unzip GMN-vul.zip

In the $VSIM_HOME directory, please run mkdir datasets;

Then, please download the datasets from the following links

Please unzip binkit_xc.zip.
Please unzip GMN-D2.zip with password ndss26ae.

Those binaries are also available from the released artifacts of PEM, GMN, and BinKit (we use the ver. 2.0 dataset).

After downloading all files, extract them in $VSIM_HOME/datasets.

If you are trying to reproduce the results in our paper, please also download the following zip files and put them in the test folder.

After the preparation step, we have a folder structure like

vSim # $VSIM_HOME
├─ ida
│ └─ ida_script.py # the IDA Python script for preprocessing
├─ src # source code of vSim
├─ scripts # some useful scripts
├─ datasets # directory for storing raw binaries
│ ├─ pem-trex # binaries for cross-optimization
│ ├─ binkit_xc # binaries for cross-compiler
│ └─ Dataset-2 # GMN-dataset-2 for cross-architecture
│
├─ test
│ ├─ cross-optimization # directory for test cross-optimization
│ │ ├─ se_cmds.txt # command lines for value extraction
│ │ ├─ fp-gen.py # script for fingerprint generation
│ │ ├─ fp-no-refine-gen.py #script for fingerprint generation without filtering
│ │ ├─ test_pool_pairwise_match.py # script for running comparison
│ │ ├─ fp-refined # fingerprints with value filtering
│ │ ├─ fp-no-refined # fingerprints without value filtering
│ │ ├─ table_1 # directory for running comparison
│ │ │ ├─ O0_pool.csv
│ │ │ ├─ O2_pool.csv
│ │ │ └─ O3_pool.csv
│ │
│ ├─ cross-compiler # directory for test cross-compiler
│ │ ├─ ... # similar to cross-optimization folder
│ ├─ GMN-dataset-2
│ ├─ GMN-vul
│ │ └─Dataset-Vulnerability # GMN vulnerability dataset

Fast run

bash AE_fast_run.sh

Those commands will print the tables of vSim's results in our paper. Tables III, IV, V are self-explanatory.

The Retrieval time of Table V is the sum of printed time cost.

For Table VI, you will see

 x86:NETGEAR_R7000
0 0.875
x86:NETGEAR_R7000
0 2;1;1;1

x86:NETGEAR_R7000 means the result is the comparison between x86 and NETGEAR_R7000

0.875 is the MRR@10, 2;1;1;1 is the ranks of comparison cases. See Table 6 of GMN's paper.

Toy example

This is a toy example for demonstration. Please run

bash AE_toy.sh

Reproduce the comparison results

Use our provided fingerprints to perform comparison.

NOTE: Rerunning the comparison will require as many CPU cores as possible.

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh

Fully reproduce the whole results, including fingerprint generation

Delete the cached fingerprints and re-generate the fingerprints

NOTE: This process is extremely time-consuming and may take several days

# Show vSim's result in Table III
bash AE_rerun_cross-optimization.sh full
# Show vSim's result in Table IV
bash AE_rerun_cross-compiler.sh full
# Show vSim's result in Table V
bash AE_rerun_GMN-D2.sh full
# Show vSim's result in Table VI
bash AE_rerun_GMN-vul.sh full

Bibtex

If this work are helpful for your research, please consider citing the following bibtex entry.

@inproceedings{wang2026vsim, title={{vSim}: Semantics-aware value extraction for efficient binary code similarity analysis}, author={Wang, Huaijin and Lin, Zhiqiang}, year={2026},
booktitle={Network and Distributed Systems Security (NDSS) Symposium}
}

About

The artifact of "vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity Analysis," NDSS 2026. You may also have a look at https://doi.org/10.5281/zenodo.17751555

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages