[Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

Description

@shaltielshmid

Is your feature request related to a problem? Please describe.

I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

Describe the solution you'd like

Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

Describe alternatives you've considered

Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

Additional context

In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

Real example:

I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

Sample python code:

fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

Sample C# code:

varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

Proposed Solution

Create a static dictionary in the BPE class, which is initialized once:

varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

Then, in the BPE.cs class in the Tokenize function here, add the following check:

if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

And then in the BPEDecoder.cs file, in the Decode function here

stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

Would be happy to compile this into a PR, if relevant.

@luisquintanilla

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

      Description

      @shaltielshmid

      Is your feature request related to a problem? Please describe.

      I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

      Describe the solution you'd like

      Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

      Describe alternatives you've considered

      Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

      Additional context

      In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

      Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

      Real example:

      I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

      Sample python code:

      fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
      print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
      print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

      Sample C# code:

      varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

      Proposed Solution

      Create a static dictionary in the BPE class, which is initialized once:

      varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

      Then, in the BPE.cs class in the Tokenize function here, add the following check:

      if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

      And then in the BPEDecoder.cs file, in the Decode function here

      stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

      Would be happy to compile this into a PR, if relevant.

      @luisquintanilla

      Metadata

      Metadata

      Assignees

      No one assigned

        Type

        No type

        Projects

        No projects

          Milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

          Description

          @shaltielshmid

          Is your feature request related to a problem? Please describe.

          I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

          Describe the solution you'd like

          Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

          Describe alternatives you've considered

          Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

          Additional context

          In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

          Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

          Real example:

          I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

          Sample python code:

          fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
          print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
          print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

          Sample C# code:

          varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

          Proposed Solution

          Create a static dictionary in the BPE class, which is initialized once:

          varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

          Then, in the BPE.cs class in the Tokenize function here, add the following check:

          if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

          And then in the BPEDecoder.cs file, in the Decode function here

          stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

          Would be happy to compile this into a PR, if relevant.

          @luisquintanilla

          Metadata

          Metadata

          Assignees

          No one assigned

            Type

            No type

            Projects

            No projects

              Milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

              Description

              @shaltielshmid

              Is your feature request related to a problem? Please describe.

              I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

              Describe the solution you'd like

              Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

              Describe alternatives you've considered

              Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

              Additional context

              In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

              Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

              Real example:

              I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

              Sample python code:

              fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
              print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
              print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

              Sample C# code:

              varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

              Proposed Solution

              Create a static dictionary in the BPE class, which is initialized once:

              varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

              Then, in the BPE.cs class in the Tokenize function here, add the following check:

              if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

              And then in the BPEDecoder.cs file, in the Decode function here

              stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

              Would be happy to compile this into a PR, if relevant.

              @luisquintanilla

              Metadata

              Metadata

              Assignees

              No one assigned

                Type

                No type

                Projects

                No projects

                  Milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

                  Description

                  @shaltielshmid

                  Is your feature request related to a problem? Please describe.

                  I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

                  Describe the solution you'd like

                  Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

                  Describe alternatives you've considered

                  Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

                  Additional context

                  In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

                  Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

                  Real example:

                  I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

                  Sample python code:

                  fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
                  print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
                  print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

                  Sample C# code:

                  varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

                  Proposed Solution

                  Create a static dictionary in the BPE class, which is initialized once:

                  varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

                  Then, in the BPE.cs class in the Tokenize function here, add the following check:

                  if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

                  And then in the BPEDecoder.cs file, in the Decode function here

                  stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

                  Would be happy to compile this into a PR, if relevant.

                  @luisquintanilla

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

                      Description

                      @shaltielshmid

                      Is your feature request related to a problem? Please describe.

                      I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

                      Describe the solution you'd like

                      Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

                      Describe alternatives you've considered

                      Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

                      Additional context

                      In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

                      Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

                      Real example:

                      I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

                      Sample python code:

                      fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
                      print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
                      print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

                      Sample C# code:

                      varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

                      Proposed Solution

                      Create a static dictionary in the BPE class, which is initialized once:

                      varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

                      Then, in the BPE.cs class in the Tokenize function here, add the following check:

                      if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

                      And then in the BPEDecoder.cs file, in the Decode function here

                      stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

                      Would be happy to compile this into a PR, if relevant.

                      @luisquintanilla

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

                          Description

                          @shaltielshmid

                          Is your feature request related to a problem? Please describe.

                          I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

                          Describe the solution you'd like

                          Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

                          Describe alternatives you've considered

                          Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

                          Additional context

                          In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

                          Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

                          Real example:

                          I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

                          Sample python code:

                          fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
                          print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
                          print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

                          Sample C# code:

                          varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

                          Proposed Solution

                          Create a static dictionary in the BPE class, which is initialized once:

                          varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

                          Then, in the BPE.cs class in the Tokenize function here, add the following check:

                          if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

                          And then in the BPEDecoder.cs file, in the Decode function here

                          stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

                          Would be happy to compile this into a PR, if relevant.

                          @luisquintanilla

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              [Tokenizers] Add support for HuggingFace BPE Tokenizer format #6901

                              Description

                              @shaltielshmid

                              Is your feature request related to a problem? Please describe.

                              I'm requesting this feature after trying to use the GPT2-style tokenizer I trained using HuggingFace in my .NET code. I had trained a model and converted the model to ONNX, but the tokenizer didn't transfer. An exact description of the problem is listed down below.

                              Describe the solution you'd like

                              Add support for a flag indicating that the tokenizer came from the HuggingFace BPE trainer, and behind the scenes handle the minor changes required.

                              Describe alternatives you've considered

                              Currently I have a class I wrote which wraps a BPE trainer and applies the adjustments before every call to the ML.NET BPE tokenizer.

                              Additional context

                              In the HuggingFace BPE code they have a dictionary bytes_to_unicode() which is list of utf-8 byte and a mapping to unicode strings. They run every byte in the string through the mapping before running the BPE encoder/decoder. Examples of where it's used can be found here and here and in other places.

                              Before the encoding, they treat the string as bytes and map all the bytes to representative unicode strings, and the same thing during after the decoding.

                              Real example:

                              I trained a BPE tokenizer using HuggingFace's tokenizers.ByteLevelBPETokenizer. The merges.txt and vocab.json can be found here: https://gist.github.com/shaltielshmid/58b7c1109639eefcd714eb6bfc3eb602.

                              Sample python code:

                              fromtransformersimportGPT2Tokenizertokenizer=GPT2Tokenizer.from_pretrained('/path/to/tokenizer')
                              print(tokenizer.encode('שלום וברכה')); // [150, 662, 426, 1396]
                              print(tokenizer.decode([150, 662, 426, 1396])); //שלוםוברכה

                              Sample C# code:

                              varbpe=newBpe("/path/to/vocab.json","/path/to/merges.txt");stringphrase="שלום וברכה";Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 0, 0, 0, 0, 0, 0, 0, 0, 0, 0stringdecoded=Bpe.Decoder.Decode(newList<int>{150,662,426,1396}.Select(id =>bpe.IdToToken(id)!));Console.WriteLine(decoded);// ש׾×ķ×ĿĠ×ķ×ijר׼×Ķ// with proposed solution from down belowphrase=newstring(Encoding.UTF8.GetBytes(phrase).Select(b =>hf_encoder[b]).ToArray());Console.WriteLine(string.Join(", ",bpe.Tokenize(phrase).Select(t =>t.Id.ToString())));// 150, 662, 426, 1396decoded=Encoding.UTF8.GetString(decoded.Select(c =>(byte)hf_decoder[c]).ToArray());Console.WriteLine(decoded);// שלום וברכה

                              Proposed Solution

                              Create a static dictionary in the BPE class, which is initialized once:

                              varhf_encoder=newDictionary<int,char>();for(intc='!';c<='~';c++)hf_encoder.Add(c,(char)c);for(intc='¡';c<='¬';c++)hf_encoder.Add(c,(char)c);for(intc='®';c<='ÿ';c++)hf_encoder.Add(c,(char)c);intn=0;for(intc=0;c<256;c++){if(!hf_encoder.ContainsKey(c))hf_encoder.Add(c,(char)(256+n++));}varhf_decoder=hf_encoder.ToDictionary(kvp =>kvp.Value, kvp =>kvp.Key);

                              Then, in the BPE.cs class in the Tokenize function here, add the following check:

                              if(_isHFFormat){sequence=newstring(Encoding.UTF8.GetBytes(sequence).Select(b =>hf_encoder[b]).ToArray())}

                              And then in the BPEDecoder.cs file, in the Decode function here

                              stringret=string.Join("",tokens);if(_suffix!=null){ret=ret.Replace(_suffix," ");}if(_isHFFormat){ret=Encoding.UTF8.GetString(ret.Select(c =>(byte)hf_decoder[c]).ToArray())}returnret;

                              Would be happy to compile this into a PR, if relevant.

                              @luisquintanilla

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions