A clean, reproducible internship project demonstrating an end-to-end supervised machine learning classification pipeline using the classic Iris dataset, feature scaling, K-Nearest Neighbors, model selection, and evaluation.
Data Classification Using AI is an internship project focused on understanding and implementing the fundamentals of supervised machine learning through a practical classification problem.
The project uses the well-known Iris flower dataset to classify flowers into three species based on four numerical measurements:
- πΏ Sepal length
- πΏ Sepal width
- πΈ Petal length
- πΈ Petal width
Rather than treating the task as a black-box prediction problem, the project demonstrates the complete machine learning workflow: data validation β preprocessing β train/test splitting β feature scaling β model selection β training β prediction β evaluation.
The implementation intentionally stays focused on the assignment requirements and avoids unnecessary frameworks, databases, APIs, or cloud services.
The main objective is to build a reliable classification model that can learn from labeled Iris flower examples and predict the species of previously unseen samples.
The project demonstrates several important machine learning concepts:
Data β Preprocessing β Training β Prediction β Evaluation
It also demonstrates why evaluation should go beyond a single accuracy number by using a confusion matrix, precision, recall, and F1 score.
The project uses K-Nearest Neighbors (KNN) as the classification algorithm.
KNN predicts the class of a new observation by examining the labels of its nearest training examples. Because KNN is distance-based, feature scaling is an important part of the pipeline.
IRIS DATASET
β
βΌ
Data Validation
β
βΌ
Stratified 80/20 Split
ββββββββ΄βββββββ
β β
TRAIN TEST
β β
βΌ β
StandardScaler β
β β
βΌ β
K Selection β
β β
βΌ β
KNN Model β
β β
ββββββββ¬βββββββ
βΌ
Predict
β
ββββββββββββ΄βββββββββββ
βΌ βΌ
Confusion Matrix Macro F1
The four Iris measurements are expressed on different numerical ranges. Standardization puts the features on a comparable scale so that distance calculations used by KNN are not dominated by a feature simply because of its numerical magnitude.
The scaler is placed inside the scikit-learn Pipeline, meaning it is fitted only using the training data. This prevents test-data leakage.
The Iris benchmark contains:
| Property | Value |
|---|---|
| Total samples | 150 |
| Classes | 3 |
| Features | 4 |
| Samples per class | 50 |
| Training samples | 120 |
| Testing samples | 30 |
setosa 50 samples
versicolor 50 samples
virginica 50 samples
The assignment demonstrates K = 5 for KNN, and the implementation supports that exact demonstration.
In addition, the project includes a small model-selection step that evaluates K = 1 through K = 15 using:
- 5-fold cross-validation
- Macro F1 scoring
- Smallest-K tie breaking
The resulting selected value is K = 5, matching the assignment demonstration.
To reproduce the exact fixed-K demonstration:
python -m src.main --no-tune-k --k 5The complete pipeline has been executed successfully.
| Metric | Result |
|---|---|
| Accuracy | 93.33% |
| Macro Precision | 94.44% |
| Macro Recall | 93.33% |
| Macro F1 | 93.27% |
| Test samples | 30 |
| Correct predictions | 28 / 30 |
| Selected K | 5 |
Predicted
Setosa Versicolor Virginica
Setosa 10 0 0
Versicolor 0 10 0
Virginica 0 2 8
The classifier correctly recognizes all 10 Setosa and 10 Versicolor test samples. Two Virginica samples are classified as Versicolor, giving an overall accuracy of 93.33%.
The repository includes an automated test suite covering the important parts of the pipeline.
pytest -qVerified result:
5 passed
Tests cover:
- Dataset schema and dimensions
- Class distribution
- Stratified 80/20 split
- Model construction
- StandardScaler + KNN pipeline
- K selection
- End-to-end classification performance
git clone https://github.com/HadeedJalani/Data-classification.git
cd Data-classificationpython -m venv .venv
.\.venv\Scripts\Activate.ps1python3 -m venv .venv
source .venv/bin/activatepython -m pip install --upgrade pip
pip install -r requirements.txtpytest -qpython -m src.mainpython -m src.main --no-tune-k --k 5Data-classification/
β
βββ π data/
β βββ iris.csv # Iris dataset
β
βββ π src/
β βββ __init__.py
β βββ config.py # Project configuration
β βββ data.py # Loading, validation and splitting
β βββ evaluate.py # Metrics and reports
β βββ main.py # Main executable pipeline
β βββ model.py # Scaling, KNN and K selection
β βββ visualize.py # Evaluation visualizations
β
βββ π tests/
β βββ test_pipeline.py # Automated tests
β
βββ π docs/
β βββ demo.py
β βββ PRESENTATION_OUTLINE.md
β βββ SUBMISSION_CHECKLIST.md
β βββ TECHNICAL_REPORT.md
β
βββ π artifacts/ # Generated locally at runtime
βββ π reports/ # Generated locally at runtime
β
βββ .gitignore
βββ requirements.txt
βββ README.md
| Technology | Purpose |
|---|---|
| Python | Core programming language |
| Pandas | Dataset loading and manipulation |
| NumPy | Numerical operations |
| Scikit-learn | Scaling, KNN, model selection and evaluation |
| Matplotlib | Confusion matrix and K-selection visualizations |
| Joblib | Model persistence |
| Pytest | Automated testing |
A fixed random_state=42 is used for the train/test split so the experiment can be reproduced consistently.
The 80/20 split uses stratification to preserve the class distribution between training and testing data.
StandardScaler is part of the scikit-learn Pipeline, so preprocessing is learned from training data rather than from the complete dataset.
The project follows the supplied internship assignment closely. Computer vision and CNNs shown in the assignment as future directions are not unnecessarily added to this implementation.
After running the pipeline, the following outputs are generated locally:
artifacts/
βββ confusion_matrix.png
βββ iris_knn_model.joblib
βββ k_selection.png
βββ metrics.json
reports/
βββ classification_report.txt
These files are intentionally ignored by Git because they are reproducible outputs rather than source code.
This repository represents the implementation of Project 2: Data Classification Using AI from an Artificial Intelligence internship assignment.
The project focuses on demonstrating practical understanding of:
- Supervised learning
- Classification
- Data preprocessing
- Feature scaling
- Train/test methodology
- K-Nearest Neighbors
- Hyperparameter selection
- Confusion matrices
- Precision, recall and F1 score
- Reproducible ML engineering
The goal is not simply to produce a prediction, but to demonstrate the reasoning and engineering workflow behind a complete machine learning solution.
This project is intended primarily for educational and internship demonstration purposes.