Korean text normalization and language preparation package for LM in Kaldi-based ASR system
-
Updated
Apr 23, 2020 - Python
Korean text normalization and language preparation package for LM in Kaldi-based ASR system
Simple-to-use scoring function for arbitrarily tokenized texts.
Subword-augmented Embedding for Cloze Reading Comprehension (COLING 2018)
johnny - a neural network graph based DEPendency Parser
A framework for generating subword vocabulary from a tensorflow dataset and building custom BERT tokenizer models.
200k-vocab SentencePiece (Unigram) tokenizer for German-primary LLMs — German/English/code, low fertility, byte-fallback, chat-template tokens. From the Auralis/Helix project.
Com la tokenització fractura la morfologia catalana i si una segmentació conscient dels morfemes recupera la geometria. Provat en 3 llengües indoeuropees (català, castellà, anglès): el català es fragmenta ~1,7× més que l'anglès; forçar el tall morfèmic recupera la composicionalitat (robust a portadora i replicat en castellà).
Subword Neural Machine Translation
To associate your repository with the subword topic, visit your repo's landing page and select "manage topics."