Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

[GraphsGPT] A Graph is Worth $K$ Words:
Euclideanizing Graph using Pure Transformer (ICML2024)

Zhangyang Gao*, Daize Dong*, Cheng Tan, Jun Xia, Bozhen Hu, Stan Z. Li

Published on The 41st International Conference on Machine Learning (ICML 2024).

arXiv

Introduction

Can we model Non-Euclidean graphs as pure language or even Euclidean vectors while retaining their inherent information? The Non-Euclidean property have posed a long term challenge in graph modeling. Despite recent graph neural networks and graph transformers efforts encoding graphs as Euclidean vectors, recovering the original graph from vectors remains a challenge. In this paper, we introduce GraphsGPT, featuring an Graph2Seq encoder that transforms Non-Euclidean graphs into learnable GraphWords in the Euclidean space, along with a GraphGPT decoder that reconstructs the original graph from GraphWords to ensure information equivalence. We pretrain GraphsGPT on 100M molecules and yield some interesting findings:

  • The pretrained Graph2Seq excels in graph representation learning, achieving state-of-the-art results on $8/9$ graph classification and regression tasks.
  • The pretrained GraphGPT serves as a strong graph generator, demonstrated by its strong ability to perform both few-shot and conditional graph generation.
  • Graph2Seq+GraphGPT enables effective graph mixup in the Euclidean space, overcoming previously known Non-Euclidean challenges.
  • The edge-centric pretraining framework GraphsGPT demonstrates its efficacy in graph domain tasks, excelling in both representation and generation.

graphsgpt.svg

Installation

To get started with GraphsGPT, please run the following commands to install the environments.

git clone git@github.com:A4Bio/GraphsGPT.git --depth=1
cd GraphsGPT
conda create --name graphsgpt python=3.12
conda activate graphsgpt
pip install -e .[dev]
pip install -r requirements.txt

Quickstart

We provide some Jupyter Notebooks in ./jupyter_notebooks, and their corresponding online Google Colaboratory Notebooks. You can run them for a quick start.

Jupyter NotebookGoogle Colaboratory
GraphsGPT Pipelineexample_pipeline.ipynbOpen In Colab
Clustering Analysisclustering.ipynbOpen In Colab
Hybridization Analysishybridization.ipynbOpen In Colab
Interpolation Analysisinterpolation.ipynbOpen In Colab

Checkpoints

The model checkpoints can be downloaded from 🤗 Transformers. We provide both the foundational pretrained models with different number of Graph Words $\mathcal{W}$ (GraphsGPT-nW), and the conditional version with one Graph Word (GraphsGPT-1W-C).

Model NameModel TypeModel Checkpoint
GraphsGPT-1WFoundation Model
GraphsGPT-2WFoundation Model
GraphsGPT-4WFoundation Model
GraphsGPT-8WFoundation Model
GraphsGPT-1W-CFinetuned Model

Representation Experiments

You should first download the configurations and data for finetuning, and put them in ./data_finetune. (We also include the finetuned checkpoints in the model_zoom.zip file for a quick test.)

To evaluate the representation performance of the Graph2Seq Encoder, please run:

bash ./scripts/representation/finetune.sh

You can also toggle the --mixup_strategy for graph mixup using Graph2Seq.

Generation Experiments

For the unconditional generation with GraphGPT Decoder, please refer to README-Generation-Uncond.md.

For the conditional generation with GraphGPT-C Decoder, please refer to README-Generation-Cond.md.

To evaluate the few-shots generation performance of GraphGPT Decoder, please run:

bash ./scripts/generation/evaluation/moses.sh
bash ./scripts/generation/evaluation/zinc250k.sh

Citation

@article{gao2024graph,
title={A Graph is Worth $K$ Words: Euclideanizing Graph using Pure Transformer},
author={Gao, Zhangyang and Dong, Daize and Tan, Cheng and Xia, Jun and Hu, Bozhen and Li, Stan Z},
journal={arXiv preprint arXiv:2402.02464},
year={2024}
}

Contact Us

If you have any questions, please contact:

About

The official implementation of the ICML'24 paper "A Graph is Worth K Words: Euclideanizing Graph using Pure Transformer".

Resources

Stars

49 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages