PyTorch Implementation of ByteDance's Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech
text-to-speechdeep-neural-networkspytorchttsspeech-synthesisgenerative-modelsemi-supervised-learningglobal-style-tokensneural-ttsnon-autoregressiveparallel-tacotronnon-aremotion-transfercross-speakerconditional-layer-normalization
-
Updated
Nov 9, 2022 - Python