Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SSL4PR: Self-supervised learning for Parkinson's Recognition

This project aims at creating a DL system for Parkinson's recognition from speech. It leverages self-supervised learning models to transfer knowledge acquired by foundational models to the task of Parkinson's recognition. The project is based on the PC-GITA dataset and the proposed model is evaluated using 10-fold cross-validation.

Table of Contents

Setup

The project is based on Python 3.11 and PyTorch 2. The following command can be used to install the dependencies:

pip install -r requirements.txt

By default, the project leverages comet_ml for logging. To use it, you need to create an account on comet.ml. Then, you need to set the following environment variables:

export COMET_API_KEY=<your_api_key>export COMET_WORKSPACE=<your_project_name>

Disable logging: To disable the logging, you can set the training.use_comet key to false in the configs/*.yaml files.

Dataset and Data Splits

To make the results reproducible and comparable with the ones reported in the paper we make available the data splits used for 10-fold cross-validation. The splits are available in the pcgita_splits folder. The data is organized as follows:

pcgita_splits/
├── TRAIN_TEST_1
│ ├── test.csv
│ └── train.csv
├── TRAIN_TEST_2
│ ├── test.csv
│ └── train.csv
├── ...
└── TRAIN_TEST_10
├── test.csv
└── train.csv

The train.csv and test.csv files contain the list of the audio files used for training and testing, respectively. The path to the audio files is stored in the following format:

/PC_GITA_ROOT_PATH/monologue/sin_normalizar/pd/AVPEPUDEA0042-Monologo-NR.wav

where PC_GITA_ROOT_PATH is the root path to the PC-GITA dataset. We also provide a python script set_root_path.py to set the root path to the PC-GITA dataset. The script can be used as follows:

python set_root_path.py --old_root_path <old_root_path> --new_root_path <new_root_path>

where <old_root_path> can be set to PC_GITA_ROOT_PATH and <new_root_path> is the new root path to your instance of the PC-GITA dataset.

Note: The splits are generated to ensure that the same speaker does not appear in both the training and testing sets and the classes are balanced across the splits.

Experiments

To train the proposed model it is first needed to install the requirements and set the root path to the PC-GITA dataset. Then, the following command can be used to train the model:

python train.py --config <config_file>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml). There are several configuration parameter that can be set in the configuration file. The one reported in the paper is configs/W_config.yaml.

The training file creates 10 models, one for each fold, and compute the metrics for each fold. The metrics are reported on terminal and logged on comet.ml (if activated). Please be sure to update the following path in the configuration file:

training:
checkpoint_path: <path_to_save_checkpoints>data:
fold_root_path: <path_to_pcgita_splits>

where <path_to_save_checkpoints> is the path where the checkpoints will be saved and <path_to_pcgita_splits> is the path to the pcgita_splits folder.

Pre-trained models are available on the Hugging Face model hub (SSL4PR WavLM Base, SSL4PR HuBERT Base). To use them, please clone the repository running the following command:

# SSL4PR WavLM Base
git clone https://huggingface.co/morenolq/SSL4PR-wavlm-base
# SSL4PR HuBERT Base
git clone https://huggingface.co/morenolq/SSL4PR-hubert-base

Ensure you have git lfs installed. Each repository contains the pre-trained models, one per fold, named fold_1.pt, fold_2.pt, ..., fold_10.pt.

Inference on extended dataset

The paper presents the results of the proposed model on an extended dataset. The extended dataset is available in the extended_dataset folder. The script infer_extended.py can be used to infer the model on the extended dataset. The script can be used as follows:

python infer_extended.py --config <config_file> --training.ext_model_path <path_to_model> --ext_root_path <path_to_extended_dataset>

where <config_file> is the path to the configuration file (e.g., configs/W_config.yaml), <path_to_model> is the path to the model checkpoint and <path_to_extended_dataset> is the path to the extended_dataset folder.

Speech Enhancement: The proposed model can be used on the extended dataset in combination with speech enhancement preprocessing. To preprocess data following the same process as in the paper, the speech_enhancement folder contains the instructions to apply, VAD, dereverberation and noise reduction to the audio files.

Citation

If you use this code, results from this project or you want to refer to the paper, please cite the following paper:

@inproceedings{laquatra24_interspeech,
title = {{Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions}},
author = {Moreno {La Quatra} and Maria Francesca Turco and Torbjørn Svendsen and Giampiero Salvi and Juan Rafael Orozco-Arroyave and Sabato Marco Siniscalchi},
year = {2024},
booktitle = {{Interspeech 2024}},
pages = {1405--1409},
doi = {10.21437/Interspeech.2024-522},
issn = {2958-1796},
}

About

This repository contains the code for the paper "Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions" submitted to the INTERSPEECH 2024 conference.

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages