Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - osalinasv/lipnet: A Keras implementation of LipNet · GitHub
Skip to content

Repository files navigation

lipnet

A Keras implementation of LipNet

This is an implementation of the spatiotemporal convolutional neural network described by Assael et al. in this article. However, this implementation only tests the unseen speakers task, the overlapped speakers task is yet to be implemented.

The best training completed yet was started the 26th of September, 2018:

TaskEpochsCERWER
Unseen speakers709.3%15.7%

Setup

Prerequisites

Go to Python's official site to download and install Python version 3.6.6. If in a Unix/Linux system, follow your package manager's instructions to install the correct version of Python, some distros might already have such version. This project has not been tested in higher Python versions and it might not work properly.

If using with TensorFlow GPU, follow TensorFlow's and NVIDIA's CUDA installation guides. This proyect was tested with TensorFlow GPU 1.10.0 and CUDA 9.0.

Installation

To install all dependencies run the following command:

pip install -r requirements.txt
Depending in your Python environment the pip command might be different.

If you do not plan to use TensorFlow or TensorFlow GPU, remember to comment out and replace the line tensorflow-gpu==1.10.0 with your Keras back-end of choice.

Usage

Preprocessing

This project was trained using the GRID corpus dataset as per the original article.

Given the following directory structure:

GRID:
├───s1
│ ├───bbaf2n.mpg
│ ├───bbaf3s.mpg
│ └───...
├───s2
│ └───...
└───...
└───...

Use the preprocesing/extract.py script to process all videos into .npy binary files if the extracted lips. By default, each file has a numpy array of shape (75, 50, 100, 3). That is 75 frames each with 100 pixels in width and 50 in height with 3 channels per pixel.

usage: extract.py [-h] -v VIDEOS_PATH -o OUTPUT_PATH [-pp PREDICTOR_PATH]
[-p PATTERN] [-fv FIRST_VIDEO] [-lv LAST_VIDEO]
optional arguments:
-h, --help show this help message and exit
-v VIDEOS_PATH, --videos-path VIDEOS_PATH
Path to videos directory
-o OUTPUT_PATH, --output-path OUTPUT_PATH
Path for the extracted frames
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file
-p PATTERN, --pattern PATTERN
(Optional) File name pattern to match
-fv FIRST_VIDEO, --first-video FIRST_VIDEO
(Optional) First video index extracted in each speaker
(inclusive)
-lv LAST_VIDEO, --last-video LAST_VIDEO
(Optional) Last video index extracted in each speaker
(exclusive)

i.e:

python preprocessing\extract.py -v GRID -o data\dataset

This results in a new directory with the preprocessed dataset:

dataset:
├───s1
├───s2
└───...

The original article excluded speakers S1, S2, S20 and S22 from the training dataset.

Training

Use the train.py script to start training a model after preprocesing your dataset. You'll also need to provide a directory containing individual align files with the expected sentence:

usage: train.py [-h] -d DATASET_PATH -a ALIGNS_PATH [-e EPOCHS] [-ic]
optional arguments:
-h, --help show this help message and exit
-d DATASET_PATH, --dataset-path DATASET_PATH
Path to the dataset root directory
-a ALIGNS_PATH, --aligns-path ALIGNS_PATH
Path to the directory containing all align files
-e EPOCHS, --epochs EPOCHS
(Optional) Number of epochs to run
-ic, --ignore-cache (Optional) Force the generator to ignore the cache
file

i.e:

python train.py -d data/dataset -a data/aligns -e 70

The training is configured to use multiprocessing with 2 workers by default.

Before starting the training, the dataset in the given directory is split by 20% for validation and 80% for training. A cache file of this split is saved inside the data directory, i.e: data/dataset.cache.

Evaluating

Use the predict.py script to analyze a video or a directory of videos with a trained model:

usage: predict.py [-h] -v VIDEO_PATH -w WEIGHTS_PATH [-pp PREDICTOR_PATH]
optional arguments:
-h, --help show this help message and exit
-v VIDEO_PATH, --video-path VIDEO_PATH
Path to video file or batch directory to analize
-w WEIGHTS_PATH, --weights-path WEIGHTS_PATH
Path to .hdf5 trained weights file
-pp PREDICTOR_PATH, --predictor-path PREDICTOR_PATH
(Optional) Path to the predictor .dat file

i.e:

python predict.py -w data/res/2018-09-26-02-30/lipnet_065_1.96.hdf5 -v data/dataset_eval

Configuration

The env.py file hosts a number of configurable variables:

Related to the videos:

  • FRAME_COUNT: The number of frames to be expected for each video
  • IMAGE_WIDTH: The width in pixels for each video frame
  • IMAGE_HEIGHT: The height in pixels for each video frame
  • IMAGE_CHANNELS: The amount of channels for each pixel (3 is RGB and 1 is greyscale)

Related to the neural net:

  • MAX_STRING: The maximum amount of characters to expect as the encoded align sentence vector
  • OUTPUT_SIZE: The maximum amount of characters to expect as the prediction output
  • BATCH_SIZE: The amount of videos to read by batch
  • VAL_SPLIT: The fraction between 0.0 and 1.0 of the videos to take as the validation set

Related to the standardization:

  • MEAN_R: Arithmetic mean of the red channel in the training set
  • MEAN_G: Arithmetic mean of the green channel in the training set
  • MEAN_B: Arithmetic mean of the blue channel in the training set
  • STD_R: Standard deviation of the red channel in the training set
  • STD_G: Standard deviation of the green channel in the training set
  • STD_B: Standard deviation of the blue channel in the training set

To-do List

  • RGB standardization: Apply per-batch zero mean standardization
  • Augmentation: Make generators also output the horizontal flip of each video
  • Statistics: Record per-epoch statistics and other useful data visualizations.
  • Documentation: Proper usage and code documentation
  • Testing: Develop unit testing

Built With

  • Python - The programming language
  • Keras - The high-level neural network API

Author

  • Omar Salinas - omarsalinas16 Developed as a bachelor's thesis @ UACJ - IIT

See also the list of contributors who participated in this project.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details

About

A Keras implementation of LipNet

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages