Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation


Topic modeling is your turf too.
Contextual topic models with representations from transformers.

DOI

Features

SOTA Transformer-based Topic Models🧭 , 🔑 KeyNMF, 💎 GMM, 🗻 Topeax, 🌀 SensTopic, Clustering Models (BERTopic and Top2Vec), Autoencoding models (ZeroShotTM and CombinedTM), FASTopic
Models for all Scenarios📈 Dynamic, 🌊 Online, 🌿 Seeded, 🌲 Hierarchical, and 📷 Multimodal topic modeling
Easy Interpretation📑 Pretty Printing, 📊 Interactive Figures, 🎨 topicwizard compatible
Topic Analysis🤖 LLM-generated names and descriptions, 👋 Manual Topic Naming
Informative Topic Descriptions🔑 Keyphrases, Noun-phrases, Lemmatization, Stemming

Basics

Open in Colab

For more details on a particular topic, you can consult our documentation page:

🏠 Build and Train Topic Models🎨 Explore, Interpret and Visualize your Models🔧 Modify and Fine-tune Topic Models
📌 Choose the Right Model for your Use-Case📈 Explore Topics Changing over Time📰 Use Phrases or Lemmas for Topic Models
🌊 Extract Topics from a Stream of Documents🌲 Find Hierarchical Order in Topics🐳 Name Topics with Large Language Models

Installation

Turftopic can be installed from PyPI.

pip install turftopic

If you intend to use CTMs, make sure to install the package with Pyro as an optional dependency.

pip install "turftopic[pyro-ppl]"

If you want to use clustering models like BERTopic or Top2Vec, install:

pip install "turftopic[umap-learn]"

Fitting a Model

Turftopic's models follow the scikit-learn API conventions, and as such they are quite easy to use if you are familiar with scikit-learn workflows.

Here's an example of how you use KeyNMF, one of our models on the 20Newsgroups dataset from scikit-learn.

If you are using a Mac, you might have to install the required SSL certificates on your system in order to be able to download the dataset.

fromsklearn.datasetsimportfetch_20newsgroupsnewsgroups=fetch_20newsgroups(
subset="all",
remove=("headers", "footers", "quotes"),
)
corpus: list[str] =newsgroups.dataprint(len(corpus)) # 18846

Turftopic also comes with interpretation tools that make it easy to display and understand your results.

fromturftopicimportKeyNMFmodel=KeyNMF(20)
document_topic_matrix=model.fit_transform(corpus)

Interpreting Models

Turftopic comes with a number of pretty printing utilities for interpreting the models.

To see the highest the most important words for each topic, use the print_topics() method.

model.print_topics()
Topic IDTop 10 Words
0armenians, armenian, armenia, turks, turkish, genocide, azerbaijan, soviet, turkey, azerbaijani
1sale, price, shipping, offer, sell, prices, interested, 00, games, selling
2christians, christian, bible, christianity, church, god, scripture, faith, jesus, sin
3encryption, chip, clipper, nsa, security, secure, privacy, encrypted, crypto, cryptography
....
# Print highest ranking documents for topic 0model.print_representative_documents(0, corpus, document_topic_matrix)
DocumentScore
Poor 'Poly'. I see you're preparing the groundwork for yet another retreat from your...0.40
Then you must be living in an alternate universe. Where were they? An Appeal to Mankind During the...0.40
It is 'Serdar', 'kocaoglan'. Just love it. Well, it could be your head wasn't screwed on just right...0.39
model.print_topic_distribution(
"I think guns should definitely banned from all public institutions, such as schools."
)
Topic nameScore
7_gun_guns_firearms_weapons0.05
17_mail_address_email_send0.00
3_encryption_chip_clipper_nsa0.00
19_baseball_pitching_pitcher_hitter0.00
11_graphics_software_program_3d0.00

Automated Topic Naming

Turftopic now allows you to automatically assign human readable names to topics using LLMs or n-gram retrieval!

You will need to pip install "turftopic[openai]" for this to work.

fromturftopicimportKeyNMFfromturftopic.analyzersimportOpenAIAnalyzermodel=KeyNMF(10).fit(corpus)
namer=OpenAIAnalyzer("gpt-4o-mini")
model.rename_topics(namer)
model.print_topics()
Topic IDTopic NameHighest Ranking
0Operating Systems and Softwarewindows, dos, os, ms, microsoft, unix, nt, memory, program, apps
1Atheism and Belief Systemsatheism, atheist, atheists, belief, religion, religious, theists, beliefs, believe, faith
2Computer Architecture and Performancemotherboard, ram, memory, cpu, bios, isa, speed, 486, bus, performance
3Storage Technologiesdisk, drive, scsi, drives, disks, floppy, ide, dos, controller, boot
...

Vectorizers Module

You can use a set of custom vectorizers for topic modeling over phrases, as well as lemmata and stems.

You will need to pip install "turftopic[spacy]" for this to work.

fromturftopicimportBERTopicfromturftopic.vectorizers.spacyimportNounPhraseCountVectorizermodel=BERTopic(
n_components=10,
vectorizer=NounPhraseCountVectorizer("en_core_web_sm"),
)
model.fit(corpus)
model.print_topics()
Topic IDHighest Ranking
...
3fanaticism, theism, fanatism, all fanatism, theists, strong theism, strong atheism, fanatics, precisely some theists, all theism
4religion foundation darwin fish bumper stickers, darwin fish, atheism, 3d plastic fish, fish symbol, atheist books, atheist organizations, negative atheism, positive atheism, atheism index
...

Visualization

Turftopic comes with a number of visualization and pretty printing utilities for specific models and specific contexts, such as hierarchical or dynamic topic modelling. You will find an overview of these in the Interpreting and Visualizing Models section of our documentation.

pip install "turftopic[datamapplot, openai]"
fromturftopicimportClusteringTopicModelfromturftopic.analyzersimportOpenAIAnalyzermodel=ClusteringTopicModel(feature_importance="centroid").fit(corpus)
namer=OpenAIAnalyzer("gpt-5-nano")
model.rename_topics(namer)
fig=model.plot_clusters_datamapplot()
fig.show()
image

In addition, Turftopic is natively supported in topicwizard, an interactive topic model visualization library, is compatible with all models from Turftopic.

pip install "turftopic[topic-wizard]"

By far the easiest way to visualize your models for interpretation is to launch the topicwizard web app.

importtopicwizardtopicwizard.visualize(corpus, model=model)

Screenshot of the topicwizard Web Application

Alternatively you can use the Figures API in topicwizard for individual HTML figures.

Citation

Please cite us when using Turftopic:

@article{
Kardos2025,
title = {Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers},
doi = {10.21105/joss.08183},
url = {https://doi.org/10.21105/joss.08183},
year = {2025},
publisher = {The Open Journal},
volume = {10},
number = {111},
pages = {8183},
author = {Kardos, Márton and Enevoldsen, Kenneth C. and Kostkan, Jan and Kristensen-McLachlan, Ross Deans and Rocca, Roberta},
journal = {Journal of Open Source Software} }

References

  • Kardos, M., Kostkan, J., Vermillet, A., Nielbo, K., Enevoldsen, K., & Rocca, R. (2024, June 13). $S^3$ - Semantic Signal separation. arXiv.org. https://arxiv.org/abs/2406.09556
  • Wu, X., Nguyen, T., Zhang, D. C., Wang, W. Y., & Luu, A. T. (2024). FASTopic: A Fast, Adaptive, Stable, and Transferable Topic Modeling Paradigm. ArXiv Preprint ArXiv:2405.17978.
  • Grootendorst, M. (2022, March 11). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv.org. https://arxiv.org/abs/2203.05794
  • Angelov, D. (2020, August 19). Top2VEC: Distributed representations of topics. arXiv.org. https://arxiv.org/abs/2008.09470
  • Bianchi, F., Terragni, S., & Hovy, D. (2020, April 8). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. arXiv.org. https://arxiv.org/abs/2004.03974
  • Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676–1683). Association for Computational Linguistics.
  • Kristensen-McLachlan, R. D., Hicke, R. M. M., Kardos, M., & Thunø, M. (2024, October 16). Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media. arXiv.org. https://arxiv.org/abs/2410.12791

About

Robust and fast topic models with sentence-transformers.

Topics

Resources

Code of conduct

Contributing

Stars

120 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages