Uh oh!
There was an error while loading. Please reload this page.
Support Gpt-5.2 in Tokenizer library - #7571
Conversation
There was a problem hiding this comment.
Pull request overview
This PR adds support for the GPT-5.2 model to the Tokenizer library, following the same pattern as GPT-5 and GPT-5.1 models which use the O200kBase encoding.
Changes:
- Added GPT-5.2 model support with O200kBase encoding
- Updated test cases to include GPT-5.2 and GPT-5.2-mini variants
- Added model mappings for both exact match and prefix matching
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
| src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs | Added "gpt-5.2" and "gpt-5.2-" mappings to O200kBase encoding in both exact match dictionary and prefix array |
| test/Microsoft.ML.Tokenizers.Tests/TiktokenTests.cs | Added GPT5_2 tokenizer property, included it in O200kBase encoding test, and added test cases for "gpt-5.2" and "gpt-5.2-mini" model names |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@## main #7571 +/- ##
=======================================
Coverage 69.02% 69.02% =======================================
Files 1482 1482 Lines 274096 274099 +3 Branches 28266 28266 =======================================
+ Hits 189191 189199 +8
Misses 77518 77518 + Partials 7387 7382 -5
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
tarekgh
commented
Jan 22, 2026
/ba-g unrelated failures |
Uh oh!
There was an error while loading. Please reload this page.
No description provided.