Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Tokenizers

Go bindings for the HuggingFace Tokenizers library.

Installation

make build to build libtokenizers.a that you need to run your application that uses bindings. In addition, you need to inform the linker where to find that static library: go run -ldflags="-extldflags '-L./path/to/libtokenizers/directory'" . or just add it to the CGO_LDFLAGS environment variable: CGO_LDFLAGS="-L./path/to/libtokenizers/directory" to avoid specifying it every time.

Using pre-built binaries

If you don't want to install Rust toolchain, build it in docker: docker build --platform=linux/amd64 -f release/Dockerfile . or use prebuilt binaries from the releases page.

Links to prebuilt libraries

Getting started

TLDR: working example.

Load a tokenizer from a JSON config:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromFile("./data/bert-base-uncased.json")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Load a tokenizer from Huggingface:

import"github.com/daulet/tokenizers"tk, err:=tokenizers.FromPretrained("google-bert/bert-base-uncased")
iferr!=nil {
returnerr
}
// release native resourcesdefertk.Close()

Encode text and decode tokens:

fmt.Println("Vocab size:", tk.VocabSize())
// Vocab size: 30522fmt.Println(tk.Encode("brown fox jumps over the lazy dog", false))
// [2829 4419 14523 2058 1996 13971 3899] [brown fox jumps over the lazy dog]fmt.Println(tk.Encode("brown fox jumps over the lazy dog", true))
// [101 2829 4419 14523 2058 1996 13971 3899 102] [[CLS] brown fox jumps over the lazy dog [SEP]]fmt.Println(tk.Decode([]uint32{2829, 4419, 14523, 2058, 1996, 13971, 3899}, true))
// brown fox jumps over the lazy dog

If you want explicit error handling for encode/decode calls, use EncodeErr, EncodeWithOptionsErr, and DecodeErr.

Encode text with options:

varencodeOptions []tokenizers.EncodeOptionencodeOptions=append(encodeOptions, tokenizers.WithReturnTypeIDs())
encodeOptions=append(encodeOptions, tokenizers.WithReturnAttentionMask())
encodeOptions=append(encodeOptions, tokenizers.WithReturnTokens())
encodeOptions=append(encodeOptions, tokenizers.WithReturnOffsets())
encodeOptions=append(encodeOptions, tokenizers.WithReturnSpecialTokensMask())
// Or just basically// encodeOptions = append(encodeOptions, tokenizers.WithReturnAllAttributes())encodingResponse:=tk.EncodeWithOptions("brown fox jumps over the lazy dog", false, encodeOptions...)
fmt.Println(encodingResponse.IDs)
// [2829 4419 14523 2058 1996 13971 3899]fmt.Println(encodingResponse.TypeIDs)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.SpecialTokensMask)
// [0 0 0 0 0 0 0]fmt.Println(encodingResponse.AttentionMask)
// [1 1 1 1 1 1 1]fmt.Println(encodingResponse.Tokens)
// [brown fox jumps over the lazy dog]fmt.Println(encodingResponse.Offsets)
// [[0 5] [6 9] [10 15] [16 20] [21 24] [25 29] [30 33]]

Benchmarks

Tiktoken vs HuggingFace

Tiktoken is 3x faster on most tasks.

> go test . -ldflags="-extldflags '-L.'" -run=^\$ -bench=. -benchmem -count=1 -benchtime=1s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
cpu: Apple M1 Pro
BenchmarkEncodeNTimes/huggingface-10 133966 10456 ns/op 256 B/op 12 allocs/op
BenchmarkEncodeNTimes/tiktoken-10 339538 3759 ns/op 88 B/op 4 allocs/op
BenchmarkEncodeNChars/huggingface-10 456006800 2.798 ns/op 0 B/op 0 allocs/op
BenchmarkEncodeNChars/tiktoken-10 615315394 2.959 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes/huggingface-10 817164 1489 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTimes/tiktoken-10 2369224 513.9 ns/op 64 B/op 2 allocs/op
BenchmarkDecodeNTokens/huggingface-10 7423770 170.8 ns/op 4 B/op 0 allocs/op
BenchmarkDecodeNTokens/tiktoken-10 80597544 19.40 ns/op 4 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 40.626s

Go vs Rust

go test . -run=^\$ -bench=. -benchmem -count=10 > test/benchmark/$(git rev-parse HEAD).txt

Decoding overhead (due to CGO and extra allocations) is between 2% to 9% depending on the benchmark.

go test. -bench=. -benchmem -benchtime=10s
goos: darwin
goarch: arm64
pkg: github.com/daulet/tokenizers
BenchmarkEncodeNTimes-10 959494 12622 ns/op 232 B/op 12 allocs/op
BenchmarkEncodeNChars-10 1000000000 2.046 ns/op 0 B/op 0 allocs/op
BenchmarkDecodeNTimes-10 2758072 4345 ns/op 96 B/op 3 allocs/op
BenchmarkDecodeNTokens-10 18689725 648.5 ns/op 7 B/op 0 allocs/op
PASS
ok github.com/daulet/tokenizers 126.681s

Run equivalent Rust tests with cargo bench.

decode_n_times time: [3.9812 µs 3.9874 µs 3.9939 µs]
change: [-0.4103% -0.1338% +0.1275%] (p = 0.33 > 0.05)
No change in performance detected.
Found 7 outliers among 100 measurements (7.00%)
7 (7.00%) high mild
decode_n_tokens time: [651.72 ns 661.73 ns 675.78 ns]
change: [+0.3504% +2.0016% +3.5507%] (p = 0.01 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
2 (2.00%) high mild
5 (5.00%) high severe

Contributing

Please refer to CONTRIBUTING.md for information on how to contribute a PR to this project.

About

Go, Wasm bindings for HF Tokenizers and Tiktoken

Topics

Resources

Contributing

Stars

225 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages