This repository combines `WavLM`, a powerful speech representation model from Microsoft, with `MSDD` (Multi-Scale Diarization Decoder), a state-of-the-art approach for speaker diarization from Nvidia.
-
Updated
Jun 17, 2025 - Jupyter Notebook
This repository combines `WavLM`, a powerful speech representation model from Microsoft, with `MSDD` (Multi-Scale Diarization Decoder), a state-of-the-art approach for speaker diarization from Nvidia.
Code accompanying our paper on finetuning self-supervised general speech representations with a combination of contrastive and non-contrastive methods.
Add a description, image, and links to the speech-embedding topic page so that developers can more easily learn about it.
To associate your repository with the speech-embedding topic, visit your repo's landing page and select "manage topics."