Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - scienceguyrob/Stuffed: A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - scienceguyrob/Stuffed: A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - scienceguyrob/Stuffed: A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - scienceguyrob/Stuffed: A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - scienceguyrob/Stuffed: A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); GitHub - scienceguyrob/Stuffed: A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation


Stuffed

Stuffed Framework - Useful for testing algorithms on unlabelled data streams.

Author: Rob Lyon, School of Computer Science & Jodrell Bank Centre for Astrophysics, University of Manchester, Kilburn Building, Oxford Road, Manchester M13 9PL.

Contact: rob@scienceguyrob.com or robert.lyon@postgrad.manchester.ac.uk Web: http://www.scienceguyrob.com or http://www.cs.manchester.ac.uk or alternatively http://www.jb.man.ac.uk


  1. Overview

    Stuffed is a wrapper for WEKA and MOA classification algorithms, which enables testing and evaluation on unlabelled data streams. This is (or was last I checked) hard to achieve with MOA. Stuffed makes this possible by using custom sampling methods to sample large data sets so that they can contain:

     - Varied levels of class balance in both test and training sets.
    - Varied levels of labelling in the test data streams.
    

    The custom sampling method produces meta data with each sampling, that allows stream classifier predictions to be evaluated on unlabelled data. For instance, if a data item in the stream is unlabelled (?), typical evaluation mechanisms would not evaluate classifier performance on this example. However since Stuffed keeps meta data at hand, it is possible to evaluate the label assigned by a classifier to each unlabelled instance.

    Stuffed is only designed to work on binary classification problems. It can be used to gather statistics on classifier performance, is easily extensible, and can be used with other tools such as MatLab.

    So far Stuffed has been used to perform experiments for two papers:

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. A Study on Classification in Imbalanced and Partially-Labelled Data Streams, in International Conference on Systems, Man, and Cybernetics (SMC), pages 1506-1511, 2013, IEEE.

    R. J. Lyon, J. M. Brooke, J. D. Knowles, B. W. Stappers. Hellinger Distance Trees for Imbalanced Streams, In 22nd International Conference on Pattern Recognition, pages 1969-1974, Stockholm, Sweden, 2014, IEEE.

    If you use Stuffed please use the citations below.

  2. Use

    The algorithm is designed to work directly with both the MOA stream test framework and WEKA. It is a wrapper API, thus is not meant to be executed as an application. Rather you incorporate it directly into your code projects, to be extended, refined and improved.

    The code comes with examples of how it can be executed which speaks for themselves. Also check the user manual for more information.

  3. Citing our work

    Please use the following citation if you make use of this algorithm:

    @inproceedings{Lyon:2014:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{Hellinger Distance Trees for Imbalanced Streams}}, booktitle = {22nd IEEE International Conference on Pattern Recognition}, series = {ICPR '14}, year = {2014}, month = {August}, pages = {1969-1974}, location = {Stockholm, Sweden}, publisher = {IEEE} }

    @inproceedings{Lyon:2013:jk, author = {{Lyon}, R.~J. and {Knowles}, J.~D. and {Brooke}, J.~M. and {Stappers}, B.~W.}, title = {{A Study on Classification in Imbalanced and Partially-Labelled Data Streams}}, booktitle = {International Conference on Systems, Man, and Cybernetics}, series = {SMC '13}, year = {2013}, month = {October}, pages = {1506-1511}, location = {Manchester, United Kingdom}, publisher = {IEEE} }

  4. Acknowledgements

    This work was supported by grant EP/I028099/1 for the University of Manchester Centre for Doctoral Training in Computer Science, from the UK Engineering and Physical Sciences Research Council (EPSRC).

About

A framework useful for evaluating static and streaming classifiers, on large and unlabelled datasets.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages