Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Visual Prompt Engineering for Multimodal and Irregularly Sampled Medical Data

Malte Tölle, Mohamad Scharaf, Samantha Fischer, Christoph Reich, Silav Zeid, Christoph Dieterich, Benjamin Meder, Norbert Frey, Philipp Wild, Sandy Engelhardt

Paper link: https://arxiv.org/abs/2501.18237

Abstract

A multitude of examinations are conducted to assess a patient's health, with each modality contributing unique information that collectively creates comprehensive understanding. These assessments include temporal data with varying sampling rates as well as single value measurements, interventions like medications, or imaging modalities. While physicians are able to process different information easily, neural networks need specific modeling for each modality complicating the training procedure. We demonstrate that this complexity can be significantly reduced by visualizing all information as images along with unstructured text and subsequently training a conventional vision-text transformer. Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), not only simplifies data preprocessing and modeling but also outperforms current state-of-the-art methods in predicting in-hospital mortality and phenotyping, as evaluated on 6,175 patients from the MIMIC-IV dataset. The modalities include patient's clinical measurements, medications, X-ray images, and electrocardiography scans. % characteristics, conditions, and We hope our work inspires advancements in multi-modal medical AI by reducing the training complexity to (visual) prompt engineering, thus lowering entry barriers and enabling no-code solutions for training.

Method

During a hospital stay, a patient typically undergoes multiple examinations, each offering distinct insights into their health status. While physicians have learned to intuitively extract the different information and assemble them to an overall picture, neural networks need specific modeling of the different modalities and their interactions. Nevertheless, once these challenges are addressed, multi-modal models have demonstrated promising performance. However, a significant challenge persists: How to integrate multi-modal data that is captured at irregularly sampled time intervals?

Description

Our primary contribution is a substantial reduction of the modeling complexity for multiple irregularly sampled modalities by transforming each modality into an image representation. Humans are then tasked with visualizing the different modalities in an informative manner, effectively engaging in a form of "visual prompt engineering". For example, laboratory measurements can be represented as line graphs over time to convey trends and patterns (Li et al., 2023). Our approach, Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM), unifies the data processing pipeline, significantly reducing modeling complexity. This approach not only mimics the way humans interpret diverse data streams but also demonstrates significant improvements across a range of tasks.

Description

Results

Results of ViTiMM for the task In-hospital Mortality and Phenotyping compared to MeTra and MedFuse.
We compare the three methods for uni-modal training for clinical measurements (C) and X-ray (X) as well as a combination of the two similar to their original publication.
Extension of both methods to further modalities requires explicit modeling, which must not be done in ViTiMM.
Thus, by only plotting the other modalities, our method can straightforwardly expand to arbitrary modalities.
The results per phenotype can be found in Supplementary Table.
The corresponding significance tests (pairwise t-test) can be found in Supplementary Tables.

Modalities:

  • C: Clinical measurements
  • X: CXR images
  • M: Medications
  • E: Electrocardiography
MethodModalitiesIn-hospital MortalityPhenotyping
AUROCAUPRCBal. Acc.AUROCAUPRCBal. Acc.
MeTraC0.7910.4410.6090.6910.4000.574
X0.8100.4710.5440.6670.3870.564
C|X0.8590.5950.7070.7120.4310.583
MedFuseC0.8120.4480.5710.7050.4170.569
X0.6620.2640.5000.6400.3490.538
C|X0.8050.4310.6310.7330.4480.600
ViTiMM (Ours)C0.8370.5120.7430.7660.5060.618
X0.8260.4940.7580.7300.4600.589
M0.7410.3460.6800.7100.4300.577
E0.7040.2970.6360.6810.4270.573
C|X0.8750.6150.7760.7780.5300.636
C|M|X|E0.9220.7640.8470.7840.5490.659

Usage

After downloading the MIMIC datasets all plots can be created with the plot_[labs,ecgs,meds].ipynb files.

MIMIC-IV: https://physionet.org/content/mimiciv/3.1/

MIMIC-CXR: https://physionet.org/content/mimic-cxr-jpg/2.1.0/

MIMIC-IC-ECG: https://physionet.org/content/mimic-iv-ecg/1.0/

Place the runs folder in this directory, the data directory can have an arbitrary location.

Training can be performed with:

python main.py \
--task [inhospital_mortality,phenotyping] \
--model [swin,vit] \
--modalities lab med cxr ecg \
[--with_text] \
--root PATH_TO_DATA \
--n_epochs 3 \
--weight_decay 3e-8 \
--lrs 1e-5 5e-6 1e-6 \
--batch_size 4 \
[--ckpt PATH_TO_CKPT] \
--seed 0

BibTeX

@misc{toelle2025vitimm,
title={Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers},
author={T{\"o}lle, Malte and Scharaf, Mohamad and Fischer, Samantha and Reich, Christoph and Zeid, Silav and Dieterich, Christoph and Meder, Benjamin and Frey, Norbert and Wild, Philipp and Engelhardt, Sandy},
year={2025},
doi={10.48550/arXiv.2501.18237}
}

Contact

Malte Tölle
malte.toelle@med.uni-heidelberg.de
@maltetoelle

Group Artificial Intelligence in Cardiovascular Medicine (AICM)
Heidelberg University Hospital
Im Neuenheimer Feld 410, 69120 Heidelberg, Germany

About

Vision Transformer for irregular sampled Multi-modal Measurements (ViTiMM)

Resources

Stars

0 stars

Watchers

1 watching

Forks

Used by

Contributors

Languages