Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - UK183/Email-Classifier-using-SVM: Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification. · GitHub
Skip to content

Repository files navigation

📧 Email Classification using SVM & Flask

An end-to-end Machine Learning web application that classifies emails as Spam or Legitimate (Ham) using a Support Vector Machine (SVM) model trained on real-world email data.
Deployed via Flask, this project demonstrates the complete ML lifecycle — from text preprocessing and feature engineering to web-based deployment — delivering an interactive, explainable, and production-ready application.


🎯 Project Outcomes

  • ✅ Achieved over 98% accuracy in classifying spam and legitimate emails.
  • ✅ Deployed an interactive Flask web app enabling real-time email spam detection.
  • ✅ Implemented TF-IDF feature extraction and SVM optimization for high-precision classification.
  • ✅ Designed an explainable interface showing influential words and weights driving predictions.
  • ✅ Enhanced user experience with a “Get Detail” feature and search filter to explore model insights.

This project reflects skills in data preprocessing, feature engineering, model evaluation, and Flask deployment, aligning directly with Data Science and Machine Learning engineering roles.


🧠 Technical Stack

  • Programming Language: Python
  • Frameworks: Flask, Scikit-learn
  • ML Algorithm: Support Vector Machine (SVM)
  • Feature Extraction: TF-IDF Vectorizer
  • Libraries: NumPy, Pandas, Joblib
  • Frontend: HTML, CSS, JavaScript

📊 Project Workflow

  1. Data Preprocessing

    • Cleaned and normalized email text (lowercasing, punctuation & digit removal).
    • Converted textual data into numerical vectors using TF-IDF.
  2. Model Building

    • Trained an SVM classifier for binary text classification.
    • Tuned parameters using GridSearchCV for optimal accuracy.
  3. Model Evaluation

    • Evaluated using confusion matrix, precision, recall, F1-score, and accuracy metrics.
    • Ensured balanced performance across both spam and ham categories.
  4. Deployment

    • Integrated the trained model with a Flask web interface.
    • Enabled real-time predictions and model interpretability features.

💡 Real-World Application

This solution can be extended to:

  • 📬 Enterprise email security systems
  • 🔎 Phishing or fraud detection platforms
  • 💬 Chat moderation and text classification tools

📊 Project Flow

Data Importing → Load dataset (spam.csv) and clean unnecessary columns
Preprocessing → Apply label encoding, remove duplicates, and perform basic text cleaning (lowercase, punctuation removal)
EDA → Visualize spam vs ham distribution and analyze message length patterns
Vectorization → Convert text data into numerical features using CountVectorizer
TF-IDF Transformation → Reweight words based on their importance and frequency
SVM Model → Train a Support Vector Machine classifier for spam detection
Evaluation → Measure model performance using accuracy, confusion matrix, and classification report
Prediction → Test the model on new email examples through an interactive Flask web app


🧩 Folder Structure

Email_Classifier_SVM/

├── app.py

├── model.pkl

├── vector.pkl

├── tf.pkl

├──index.html

├── style.css

└── requirements.txt


⚙️ How to Run

git clone https://github.com/<your-username>/Email-Classifier-using-SVM.git
cd Email-Classifier-using-SVM
python -m venv venv
venv\Scripts\activate # Windows# ORsource venv/bin/activate # macOS/Linux
pip install -r requirements.txt
python app.py

Then open your browser at 👉 http://127.0.0.1:5000

📈 Example Predictions

📧 Input Email🧠 Predicted Output
"Congratulations! You’ve won a free iPhone. Click here to claim now!"🚫 Spam Email
"Team meeting scheduled at 10 AM tomorrow."Legitimate (Ham)
"Get 50% off on all products! Limited time offer."🚫 Spam Email
"Your invoice for the last month is attached below."Legitimate (Ham)
"Win cash rewards by completing this short survey!"🚫 Spam Email

🏆 Key Achievements & Skills Demonstrated

  • End-to-End ML Pipeline: Data preprocessing → model training → deployment
  • Text Analytics:Text Preprocessing & Feature extraction via CountVectorizer & TF-IDF
  • Model Optimization: Hyperparameter tuning with GridSearchCV
  • Web Deployment: Flask integration and UI development
  • Explainable AI (XAI): Display of top influential words for transparency
  • Full-Stack ML Project Execution: From dataset to live application

👤 Author

Kazi Umar
Linkedin profile: https://www.linkedin.com/in/umar-kazi18
💼 Data Analyst | ML Engineer | Data Science & AI Enthusiast | Power BI | Python | SQL

About

Built a real-world email spam classifier using Support Vector Machine(SVM), achieving 98% accuracy through robust text preprocessing, TF-IDF feature extraction, and EDA. Deployed the model with Flask, enabling real-time predictions and visualization of words influencing classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages