Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - deeplearningplus/tGPT: Generative Pretraining from Transcriptomes · GitHub
Skip to content

Repository files navigation

tGPT:Generative pretraining from large-scale transcriptomes

tGPT

"Generative pretraining from large-scale transcriptomes: Implications for single-cell deciphering and clinical translation".

Introduction of tGPT

Exponential accumulation of single-cell transcriptomes poses great challenge for efficient assimilation. Here, we present an approach entitled tGPT towards integration of 22.3 million single-cell transcriptomes by modeling gene expression rankings as generative pretraining task tGPT is conceptually simple in that it autoregressively models the ranking of a gene in the context of its preceding neighbors. We demonstrated the high performance of tGPT on a range of fundamental single-cell analysis tasks and novel applications on bulk tissues. The single-cell clusters and cell lineage trajectories derived from tGPT are highly aligned with known cell labels and states. The feature patterns of tumor bulk tissues learned by tGPT are associated with a wide range of genomic alteration events, prognosis and treatment outcome of immunotherapy. tGPT represents a new analytical paradigm for integrating and deciphering massive amount of transcriptome data and it will facilitate the interpretation and clinical translation of single-cell transcriptomes

Architecture of tGPT

image

This figure consists of three components: development of tGPT, applications of tGPT for single-cell and bulk tissue transcriptomes.

The following packages are required:

The following example was tested with:

python==3.7.11 numpy==1.20.0
torch==1.7.1
scanpy==1.9.1

Using tGPT on own data

Load packages

importreimportosimportsysimportgzipimporttorchimportnumpyasnpimportpandasaspdimportscanpyasscfromtqdmimporttqdmfromtorch.utils.dataimportDataLoader, DatasetfromtransformersimportPreTrainedTokenizerFast, GPT2LMHeadModel, GPT2Model

Setting parameter and file path

device="cuda"iftorch.cuda.is_available() else"cpu"tokenizer_file="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 62 top expressing genes.checkpoint="lixiangchun/transcriptome-gpt-1024-8-16-64"## Pretrained model with sequence of 126 top expressing genes.#checkpoint = "lixiangchun/transcriptome-gpt-1024-8-16-64"celltype_path="./data/Muris_cell_labels.txt.gz"## Cell type annotationmax_len=64## Number of top genes used for analysistext_file="./data/Muris_gene_rankings.txt.gz"## Gene symbols ranked by exprssion

Extract features

classLineDataset(Dataset):
def__init__(self, lines):
self.lines=linesself.regex=re.compile(r'\-|\.')
def__getitem__(self, i):
returnself.regex.sub('_', self.lines[i])
def__len__(self):
returnlen(self.lines)
tokenizer=PreTrainedTokenizerFast.from_pretrained(tokenizer_file)
model=GPT2LMHeadModel.from_pretrained(checkpoint,output_hidden_states=True).transformermodel=model.to(device)
model.eval()
lines= [s.decode().strip() forsingzip.open(text_file, "r").readlines()]
ds=LineDataset(lines)
dl=DataLoader(ds, batch_size=64)
Xs= []
foraintqdm(dl, total=len(dl)):
batch=tokenizer(a, max_length=max_len, truncation=True, padding=True, return_tensors="pt")
fork, vinbatch.items():
batch[k] =v.to(device)
withtorch.no_grad():
x=model(**batch)
eos_idxs=batch.attention_mask.sum(dim=1) -1xx=x.last_hidden_stateresult_list= [[] foriinrange(len(xx))]
forj, iteminenumerate(xx):
result_list[j] =item[1:int(eos_idxs[j]),:].mean(dim=0).tolist()
Xs.extend(result_list)
features=np.stack(Xs)

Visualization

adata=sc.AnnData(features)
celltype=pd.read_csv(celltype_path, header=None)[0].tolist()
adata.obs["celltype"] =celltypeadata.obs["celltype"] =adata.obs["celltype"].astype("category")
sc.pp.neighbors(adata,n_neighbors=20)
sc.tl.leiden(adata,resolution=0.6)
sc.tl.umap(adata)
################# Cell Type #######################sc.pl.umap(adata, color= ["celltype"], show=True)
############ Single-cell Clustering #############sc.pl.umap(adata, color= ["leiden"], show=True)

imageimage

About

Generative Pretraining from Transcriptomes

Resources

Stars

17 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages