Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🐦 Disaster Tweets — NLP Binary Classification

PythonNLPscikit-learnNLTKLicenseStatus

This repository contains an end-to-end NLP classification project developed as part of a MasterSchool Data Science program.

The objective is to build a feature-driven pipeline to classify tweets as disaster-related or not, combining interpretable linguistic feature engineering with TF-IDF vectorization and linear classification models.


📌 Project Overview

Twitter is a major channel for real-time reporting of emergency events. The ability to automatically identify genuine disaster signals from informal, figurative, or ambiguous language has direct applications in early warning systems and crisis monitoring.

The classification task is non-trivial: disaster-related vocabulary frequently appears in non-disaster contexts, creating semantic ambiguity that surface-level features cannot fully resolve.

This project:

  • performs exploratory data analysis to identify discriminative linguistic patterns,
  • engineers interpretable features from both raw and cleaned text,
  • builds a preprocessing pipeline combining word-level TF-IDF, character-level TF-IDF, and scaled numeric features,
  • trains and tunes two linear classifiers with cross-validation,
  • and evaluates model behavior with attention to the asymmetric cost of false negatives.

🎯 Problem Statement

The core challenge is semantic ambiguity.

The same words, fire, flood, crash, bomb, appear in both real emergencies and everyday conversation. A tweet saying "she's a suicide bomb" uses the same vocabulary as a tweet reporting a real attack. No word-frequency-based model can fully resolve this overlap.

This motivates a feature engineering strategy that goes beyond raw word counts, incorporating grammatical signals, composite linguistic scores, and semantic distinctiveness measures.


🧭 Methodology

The project follows a structured analytical pipeline:

Exploratory Data Analysis Surface-level linguistic signals, lexical features, POS distributions, and n-gram analysis — identifying discriminative patterns before any modeling decision.

Text Preprocessing Noise removal (URLs, mentions, hashtags), character normalization, and lemmatization — designed to preserve signal while eliminating platform-specific artifacts.

Feature Engineering

  • Surface signals extracted from raw text: num_numbers, has_link, num_exclam, num_question, has_ellipsis
  • POS count features from cleaned text: past/present verbs, common nouns, pronouns
  • Weighted composite scores: event_score (factual, event-oriented language) and social_score (conversational, informal language)
  • Semantic indices: ambiguity_ratio and distinctive_signal, built from lemmatized nouns, verbs, and pronouns on the training set only

Vectorization

  • Word-level TF-IDF with ngram_range=(1,2) captures compositional lexical patterns
  • Character-level TF-IDF with ngram_range=(3,4) captures subword patterns and spelling variation
  • StandardScaler applied to numeric

All preprocessing steps are combined in a single ColumnTransformer and wrapped in a sklearn Pipeline together with the classifier. This ensures that vectorization and scaling are fitted exclusively on the training set, with no leakage into the test set.

Modeling

  • Logistic Regression and Linear SVM
  • Both tuned with GridSearchCV (5-fold CV, optimizing F1 on the disaster class)
  • Threshold analysis on Logistic Regression probability estimates

🧠 Feature Engineering Design

Feature engineering is grounded in the exploratory analysis, not applied generically.

Event Score aggregates signals associated with factual, event-oriented language: number presence, link presence, past-tense verbs (VBD, VBN), and common nouns (NN). Each signal is weighted by its observed discriminative ratio (disaster/non-disaster mean) on the training set.

Social Score aggregates signals associated with conversational language: exclamation marks, question marks, present-tense verbs (VBP, VB), and pronouns (PRP, PRP$). Weights follow the same empirical approach.

Ambiguity Ratio measures the proportion of lemmas in a tweet that appear in both classes. A high value indicates vocabulary shared across disaster and non-disaster tweets, making classification harder.

Distinctive Signal counts lemmas that appear in only one class. A high value indicates vocabulary strongly associated with a specific class.

Both indices are constructed exclusively from the training set to prevent data leakage.


✅ Results

ModelClassPrecisionRecallF1Accuracy
Logistic RegressionNon-Disaster (0)0.820.870.840.82
Logistic RegressionDisaster (1)0.810.740.780.82
Linear SVMNon-Disaster (0)0.810.870.840.81
Linear SVMDisaster (1)0.810.730.770.81

Logistic Regression is selected as the final model. Beyond marginal performance advantages, it produces probability estimates enabling threshold adjustment, useful when minimizing false negatives is a priority.


⚠️ Limitations

The pipeline works with word frequencies and grammatical patterns, it does not understand meaning or context. Three categories of errors were identified:

  • Figurative language — disaster vocabulary used metaphorically
  • Noisy labels — tweets ambiguous even for human annotators
  • Vocabulary gaps — real incidents reported with place names unseen during training

These limitations motivate contextual models like BERT as a natural next step.


📊 Dataset

The dataset (train.csv) is publicly available on Kaggle: NLP with Disaster Tweets

7,613 labeled tweets: text and target (1 = disaster, 0 = non-disaster).


📁 Repository Structure

DisasterTweets/
│
├── notebooks/
│ └── disaster_tweets_NLP_classification_project.ipynb # Full analytical pipeline
│
├── reports/
│ └── presentation.pdf # Project presentation slides
│
├── .gitignore
├── requirements.txt
└── README.md

📦 Requirements

pip install -r requirements.txt

Key dependencies: pandas, numpy, scikit-learn, nltk, matplotlib, seaborn, wordcloud


👤 Author

Maria Petralia Data Science Program — MasterSchool March 2026

About

NLP binary classification pipeline to detect disaster-related tweets - feature engineering, TF-IDF vectorization, and linear models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages