Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

DFC

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

The model DFC proposed in this paper is an effective topic model for dataless text classification and filtering task. Feel free to contact me if you find any problem in the package. sqchen@whu.edu.cn

Requirements:

  • Jre1.7
  • Lucene 5.2.1

Description

Training/Testing Data Format

Each training sample is a document,which was preprocessed to remove stop words, and each word is separated by space. Example: training testing data

Catelog File

Catalog file is uesd to describe category information. The catalog file of data set '20 news groups' sholud be wrote in this form :
Example: cate
comp.graphics
cate
sci.med ...

If we want to combine several categories into one , take 'comp' and 'sci' as examples , then the file should be wrote like this :
Example:
cate
comp.graphics
comp.os.ms-windows.misc
comp.sys.ibm.pc.hardware
comp.windows.x
comp.sys.mac.hardware cate
sci.crypt
sci.med
sci.space
sci.electronics ...

Note:

The directory “catalog-classification” is for classification task without filtering.The file only contains the categories for classification task.

The directory “catalog-classificationWithFiltering” is for classification with filtering task.The file contains all the categories in dataset.The specifed categories are in the heading, the remaining are irrelevant categories.

Example: Classification with filtering task: med-space. med and space are specifed categories, others are irrelevant categories.
cate
med
cate
space
cate
sci.crypt
cate
sci.electronics ...

Seed Word File

Each line in seed word file corresponds to a category.Seed word is separated by space.Make sure that the category order in catalog file is the same as the order in seed word file.

LDA File

we take the LDA results after running LDA over 100 iterations, using the default parameter setting(Toolkit:JGibbLDA http://jgibblda.sourceforge.net/)

Parameter Setting

  • method:0:classification, 1:classification with filtering
  • relMethod:the method to estimating category word probability, 1:doc-rel, 0:topic-rel
  • iCateNum:the total number of specifed categories/relevant-topics,the variable R in paper
  • bCateNum:the setting total number of irrelevant-topics,the variable T in paper
  • btruth:the total number of irrelevant-topics, which is only used for evaluation and load documents
  • cateNum:the setting total number of categories
  • ctruth:the total number of categories, which is only used for
  • fakeSeedNum:the number of pseudo seed words
  • topicNum:the total number of general-topics, the variable B in paper
  • LDAtopicnum:the number of LDA hidden topic
  • alpha0:the variable alpha0 in paper
  • alpha1:the variable alpha1 in paper
  • alpha2:the variable alpha1 in paper
  • beta0:the variable beta0 in paper
  • beta1:the variable beta1 in paper
  • rho:the variable rho in paper
  • inter:the number of iterations
  • report:the method of evaluation, 1:macro F1 , 0:accuracy
  • maxDocLen:the max length of documents in the dataset
  • kd:the prior likelihood for document d being a relevant document, the variable k_d in paper
  • testSetPath:test set path
  • trainSetPath:train set path
  • catalogPath:category file path
  • seedwordPath:seed word file path
  • luceneIndexPath:LuceneIndex path
  • LDAwordmapPath:LDA word map path
  • LDAtassignPath:LDA model path
  • LDAphiPath:LDA phi path
  • LDAtwordPath:LDA top words path
  • resultWriter:output the predict result in file

launch the program

The main java entry is in class DfcMain.java.To launch the program there are several parameters must be setting as described above.

If you run the task of classification, you need to set the parameters bCateNum and btruth to be zero, then set the catalogPath

About

The java implementation of "Seed-Guided Topic Model for Document Filtering and Classification ", TOIS 2018.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages