Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

bookbot

a project that reads the given file and uses a neural network to generate text that looks like from the book.

the built-in neural network is MLP, Wavenet-inspired Hierarchical MLP, and a GPT network that's built with pure Pytorch (from scratch) along with batch normalization layer and Kaiming initialization.

Thanks to Andrej Karpathy for his great course on deep learning.

Available file types as of the moment:

  • PDF
  • TXT

Usage

Installation

You can try the project out by cloning the git repository

git clone https://github.com/alperiox/bookbot.git

Then just install the poetry environment and move on to the next steps.

How to train the network?

Simply run the main.py by setting up the arguments below.

You can start the training using the script like in the following:

python main.py --file=romeo-and-juliet.txt --model gpt --max_steps 100

Or if you want to have more control over the whole training, consider using a more detailed configuration:

ArgumentDefault ValueDescription
train_ratio0.8Ratio of the input data that will be used for training
file-Path to the PDF/TXT file
n_embed15Embedding vector's dimension
n_hidden400Hidden layer's dimensions (the hidden layers will be defined as n_hidden x n_hidden)
block_size10Block size to set up the dataset, it's our context window in this project
batch_size32The amount of samples that'll be processed in one go
epochs10Number of epochs to train the model
lr0.001Learning rate to update the weights
generateFalseTo run the generation mode, it's required to generate text using the pre-trained model. So you should train a model first
max_new_tokens100The amount of tokens that will be generated if generate flag is active
modelgptHierarchical mlp (hmlp), mlp model (mlp) or gpt (gpt) model to train
n_consecutive2The amount of consecutive tokens to concatenate in the hierarchical model
n_layers4Number of processor blocks in the model, check out the models in layers.py for more information about its usage
num_heads3Number of self-attention heads in the multi-head self-attention layer in GPT implementation
num_blocks2Number of layer blocks given the model. Sequential linear blocks for MLP and Hierarchical MLP, DecoderTransformerBlocks for GPT
contextNoneThe context for the text generation, please try to use a longer context than the block_size (required if generate is True)
devicecpuThe device to train the models on, available values are mps, cpu and cuda.

The training will generate several artifacts and will save them in the artifacts directory. The saved artifacts include the model, the data loaders, calculated losses along the training, and finally the tokenizer to use the constructed character-level vocabulary.

How to generate new text?

You can generate text after training a model first. That's because the generation pipeline makes use of the saved artifacts. In order to start the generation, you need to pass the generate flag:

python main.py --generate --context="Juliet," --max_new_tokens=100
>>> juliet, and have know lie thee why!

The generation will run until the wanted character length is matched.

Further plans

  • Implement debugging tools to analyze the neural network's training performance. (more like useful graphs and statistics.)
    • graphs to check out the layer outputs' distributions. layer output distributions (with extra information about mean, std and the distribution plot)
    • graphs to check the gradient flow
      • layer gradient means
      • layer gradient stds
      • ratio of amount of change in the parameters given the weights we multiply the learning rate with the layer's gradient's std and divide it by parameters' std. this ratio will be higher if gradient std is larger (grads vary too much from the mean) and the params are smaller in comparison.
      • layer grad distributions (with extra information about mean, std and the distribution plot)
      • ratio of the gradient of a specific layer to its input so if the ratio is too high, it means that the gradients are too high wtr to the input and we actually want constant but smaller updates throughout the network to not miss any local minimas etc
      • ratio of the amount of change vs the weights, the stats should be saved in L7 here
    • summary for the training
  • More modeling options such as LSTMs, RNNs, and Transformer-based architectures.
    • Wavenet? (implemented the hierarchical architecture)
    • GPT
    • GPT-2
  • GPT tokenizer implementation to further improve the generation quality.

Contributing

While I'm open to new feature ideas and stuff, please let me do the coding part since I'm trying to improve my overall understanding. Thus, I'd love to accept any feature requests as new PRs. You can reach me from Discord (@alperiox)

About

A toy project for my generative AI studies on text data. Train generative models with given book/text files with just a single script.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages