Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data (OncoChat)

替代文字

Introduction

Recent advances in artificial intelligence (AI) have highlighted the potential of large language models (LLMs) to revolutionize medical research. Models based on architectures such as Generative Pre-Training (GPT) and Bidirectional Encoder Representations from Transformers (BERT) have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records (EHRs), and assisting with diagnosis. Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.

We introduce OncoChat, a novel diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 solid tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

System requirements

This example was tested with the following environment. However, it should work on the other platforms.

Installation guide

  • Following instruction from miniconda to install Python.
  • Use the following command to install required packages.
# Install with GPU support. Check https://pytorch.org for more information. #+The following cmd install PyTorch compiled with cuda 118. 
pip install torch --index-url https://download.pytorch.org/whl/cu118
# If GPU not available, install the PyTorch compiled for CPU.
pip install torch --index-url https://download.pytorch.org/whl/cpu
# Install transformers, tokenizers and prettytable
pip install transformers==4.41.2 tokenizers==0.19.1 prettytable
  • The installation process will take about an hour. This heavily depends on your network bandwidth.

Demo

  • Step1 : Clone OncoChat locally from Github.
git clone https://github.com/deeplearningplus/OncoChat.git
  • Step 2 : Instruction Fine-tune a LLM:
sh train-oncochat-mamba-130m.sh

To execute this step, set the following parameters:
(1)--model_name_or_path : LLM checkpoints (LLM used in this study)
(2)--data_path : Training data (e.g.,data/CKP-train.json)
(3) --output_dir : Directory to save the fine-tuned model

OncoChat is composed of nine different LLMs. The remaining eight models need to be trained following the same steps and used for prediction. The demo training files and prediction results are displayed in the data folder.

  • Step3 : Get predictions from the fine-tuned model
sh predict_mamba-130m.sh

Set the following parameters for this step:
(1)--model_name_or_path : Path to the fine-tuned model
(2)--data_path : Test data (e.g.,data/CKP-test.jsonordata/CUP.json)
(3) --output_file : Path to save prediction results

Fine-tuned models are available for download from BaiduDisk (Password:1234). You can load these checkpoints for predictions instead of fine-tuning them yourself.

Dataset used in this study

The full dataset for fine-tune OncoChat is available at GENIE.

Benchmark Methods:

  • GDD-ENS:Darmofal M, Suman S, Atwal G, et al. Deep-learning model for tumor-type prediction using targeted clinical genomic sequencing data[J]. Cancer discovery, 2024, 14(6): 1064-1081.
  • OncoNPC:Moon I, LoPiccolo J, Baca S C, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary[J]. Nature medicine, 2023, 29(8): 2057-2067.

LLM used in this study

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages