Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🌸 Data Classification Using AI

Iris Flower Classification with Machine Learning

A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.


PythonScikit--learnTestsStatus


πŸ“Œ About the Project

Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.

The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:

  • 🌿 Sepal length
  • 🌿 Sepal width
  • 🌸 Petal length
  • 🌸 Petal width

Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β†’ preprocessing β†’ train/test splitting β†’ feature scaling β†’ model selection β†’ training β†’ prediction β†’ evaluation.

The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.


🎯 Project Objective

The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.

The project demonstrates several important machine learning concepts:

Data β†’ Preprocessing β†’ Training β†’ Prediction β†’ Evaluation

It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.


🧠 Machine Learning Approach

The project uses K-Nearest Neighbors (KNN) as the classification algorithm.

KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.

Model pipeline

 IRIS DATASET
β”‚
β–Ό
Data Validation
β”‚
β–Ό
Stratified 80/20 Split
β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
TRAIN TEST
β”‚ β”‚
β–Ό β”‚
StandardScaler β”‚
β”‚ β”‚
β–Ό β”‚
K Selection β”‚
β”‚ β”‚
β–Ό β”‚
KNN Model β”‚
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Predict
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό
Confusion Matrix Macro F1

Why StandardScaler?

The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.

The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.


πŸ“Š Dataset

The Iris benchmark contains:

PropertyValue
Total samples150
Classes3
Features4
Samples per class50
Training samples120
Testing samples30

Target classes

setosa 50 samples
versicolor 50 samples
virginica 50 samples

βš™οΈ Model Selection

The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.

In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:

  • 5-fold cross-validation
  • Macro F1 scoring
  • Smallest-K tie breaking

The resulting selected value is K = 5, matching the assignment demonstration.

To reproduce the exact fixed-K demonstration:

python -m src.main --no-tune-k --k 5

πŸ“ˆ Verified Results

The complete pipeline has been executed successfully.

MetricResult
Accuracy93.33%
Macro Precision94.44%
Macro Recall93.33%
Macro F193.27%
Test samples30
Correct predictions28 / 30
Selected K5

Confusion Matrix

 Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8

The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.


πŸ§ͺ Testing

The repository includes an automated test suite covering the important parts of the pipeline.

pytest -q

Verified result:

5 passed

Tests cover:

  • Dataset schema and dimensions
  • Class distribution
  • Stratified 80/20 split
  • Model construction
  • StandardScaler + KNN pipeline
  • K selection
  • End-to-end classification performance

πŸš€ Getting Started

1. Clone the repository

git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classification

2. Create a virtual environment

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1

macOS / Linux

python3 -m venv .venv
source .venv/bin/activate

3. Install dependencies

python -m pip install --upgrade pip
pip install -r requirements.txt

4. Run the tests

pytest -q

5. Run the complete pipeline

python -m src.main

6. Run the exact K=5 assignment demonstration

python -m src.main --no-tune-k --k 5

πŸ“ Project Structure

Data-classification/
β”‚
β”œβ”€β”€ πŸ“‚ data/
β”‚ └── iris.csv # Iris dataset
β”‚
β”œβ”€β”€ πŸ“‚ src/
β”‚ β”œβ”€β”€ __init__.py
β”‚ β”œβ”€β”€ config.py # Project configuration
β”‚ β”œβ”€β”€ data.py # Loading, validation and splitting
β”‚ β”œβ”€β”€ evaluate.py # Metrics and reports
β”‚ β”œβ”€β”€ main.py # Main executable pipeline
β”‚ β”œβ”€β”€ model.py # Scaling, KNN and K selection
β”‚ └── visualize.py # Evaluation visualizations
β”‚
β”œβ”€β”€ πŸ“‚ tests/
β”‚ └── test_pipeline.py # Automated tests
β”‚
β”œβ”€β”€ πŸ“‚ docs/
β”‚ β”œβ”€β”€ demo.py
β”‚ β”œβ”€β”€ PRESENTATION_OUTLINE.md
β”‚ β”œβ”€β”€ SUBMISSION_CHECKLIST.md
β”‚ └── TECHNICAL_REPORT.md
β”‚
β”œβ”€β”€ πŸ“‚ artifacts/ # Generated locally at runtime
β”œβ”€β”€ πŸ“‚ reports/ # Generated locally at runtime
β”‚
β”œβ”€β”€ .gitignore
β”œβ”€β”€ requirements.txt
└── README.md

πŸ› οΈ Technology Stack

TechnologyPurpose
PythonCore programming language
PandasDataset loading and manipulation
NumPyNumerical operations
Scikit-learnScaling, KNN, model selection and evaluation
MatplotlibConfusion matrix and K-selection visualizations
JoblibModel persistence
PytestAutomated testing

πŸ” Key Engineering Decisions

Reproducibility

A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.

Stratification

The 80/20 split uses stratification to preserve the class distribution between training and testing data.

Leakage prevention

StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.

Practical scope

The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.


πŸ“¦ Generated Artifacts

After running the pipeline, the following outputs are generated locally:

artifacts/
β”œβ”€β”€ confusion_matrix.png
β”œβ”€β”€ iris_knn_model.joblib
β”œβ”€β”€ k_selection.png
└── metrics.json
reports/
└── classification_report.txt

These files are intentionally ignored by Git because they are reproducible outputs rather than source code.


πŸŽ“ Internship Project Context

This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.

The project focuses on demonstrating practical understanding of:

  • Supervised learning
  • Classification
  • Data preprocessing
  • Feature scaling
  • Train/test methodology
  • K-Nearest Neighbors
  • Hyperparameter selection
  • Confusion matrices
  • Precision, recall and F1 score
  • Reproducible ML engineering

The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.


πŸ“œ License

This project is intended primarily for educational and internship demonstration purposes.


πŸ‘¨β€πŸ’» Author

Hadeed Jalani

Artificial Intelligence / Machine Learning Project

⭐ If you found this project useful, consider giving the repository a star.

About

Iris flower classification using Scikit-learn KNN, StandardScaler, model selection, and comprehensive ML evaluation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages