Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - KHU-MASLAB/TimeSeriesSeq2Seq: Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch · GitHub
Skip to content

Latest commit

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TimeSeriesSeq2Seq

Seq2Seq, Seq2Point modeling implementations using 1D convolution, LSTM, Attention mechanisms, Transformer, and Temporal Fusion Transformer(TFT).
The repo implements the following:

  • Basic convolution and LSTM layers implementation
  • Bahdanau attention LSTM Encoder-Decoder network by Bahdanau et al.(2014)
  • Vanilla Transformer by Vaswani et al.(2017). See architectures.transformer.Transformer.
  • Temporal Fusion Transformer by Lim et al.(2020). See architectures.tft.TemporalFusionTransformer.

Transformer-based classes always produce sequence-to-sequence outputs.
RNN-based classes can selectively produce sequence or point outputs:

  • Difference between rnn_seq2seq and rnn_seq2point is the decoder part. The former uses autoregressive LSTM decoder to generate sequence of vectors, while the latter uses MLP decoder to generate a single vector.

All network parameters are initialized from $\mathcal{N}\sim(0,0.01^2)$, except for bias initialized from torch.zeros. See architectures.init.

See Tutorial.ipynb for details.

Configuration

Supports $(B,L_{in},C_{in})\xrightarrow{network}(B,L_{out},C_{out})$ operations, where

$$\begin{aligned} B&=\text{batch\_size}\\\ L_{in}&=\text{input\_sequence\_length (variable)}\\\ C_{in}&=\text{input\_feature\_size}\\\ L_{out}&=\text{output\_sequence\_length (variable)}\\\ C_{out}&=\text{output\_feature\_size}\\\ \end{aligned}$$
  • hidden_size Hidden state size of LSTM encoder.
  • num_layers Number of stacks in CNN, LSTM encoder, LSTM decoder, and FC layers.
  • bidirectional Whether to use bidirectional LSTM encoder.
  • dropout Dropout rate. Applies to:
    Residual drop path in 1DCNN
    hidden state dropout in LSTM encoder/decoder(for every time step).
    Unlike torch.nn.LSTM, dropout is applied from the first LSTM layer.
  • layernorm Layer normalization in LSTM encoder and decoder.
  • attention Attention in LSTM decoder.
    Supports 'bahdanau' for Bahdanau style, 'dotproduct' for Dot Product style, and 'none for non-attended decoder.

Creating model instances

from architectures.rnn_seq2seq import *
from architectures.rnn_seq2point import *
from architectures.transformer import *
from architectures.tft import *
# LSTM encoder - LSTM decoder - MLP
seq2seq_lstm = LSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
) # 1DCNN+LSTM encoder - LSTM decoder
seq2seq_cnnlstm = CNNLSTMSeq2Seq(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm, attention
)
# LSTM encoder - MLP
seq2point_lstm = LSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# 1DCNN+LSTM encoder - MLP
seq2point_cnnlstm = CNNLSTMSeq2Point(
Cin, Cout, hidden_size, num_layers, bidirectional, dropout, layernorm
)
# Transformer
seq2seq_transformer = Transformer(
Cin, Cout, num_layers, n_heads, d_model, dropout, d_ff
)
# Temporal Fusion Transformer
seq2seq_tft = TemporalFusionTransformer(
Cin, Cout, num_layers, n_heads, d_model, dropout
)

Autoregressive forward operation

  • x Input to the network. Supports $(B,L_{in},C_{in})$ only.
  • y Output label for teacher forcing. Supports $(B,*,C_{out})$ only. Defaults to None (fully autoregressive).
  • teacher_forcing Teacher forcing ratio $\in [0,1]$. Defaults to -1 (fully autoregressive).
  • trg_len Target sequence length to generate. Defaults to 1.

If only x and trg_len is given as arguments, the model will autoregressively produce trg_len length of outputs.

Accessing model properties

By inheriting architectures.skeleton.Skeleton, model properties are automatically saved to attributes:

  • Parameters can be counted by model.count_params()
  • Properties are accessed using model.model_info attribute.
  • The identical model instance can be created by ModelClass(**model.model_init_args).
seq2seq_lstm.count_params()
model_info = seq2seq_lstm.model_info
model_init_args = seq2seq_lstm.model_init_args
print(model_info)
another_model_instance = LSTMSeq2Seq(**model_init_args)
Number of trainable parameters: 10422835
{'attention': 'bahdanau', 'bidirectional': True, 'dropout': 0.3, 'hidden_size': 256, 'input_size': 6, 'layernorm': True, 'num_layers': 3, 'output_size': 50}

About

Sequence-to-sequence model implementations including RNN, CNN, Attention, and Transformers using PyTorch

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages