Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Stock Price Predictor

A full-stack educational machine-learning project for predicting the next trading day's stock closing price. The project trains regression models on historical market data, exposes predictions through a Flask API, and provides a simple browser-based client for entering ticker symbols and viewing predicted price movement.

Disclaimer
This project is intended for learning, experimentation, and portfolio demonstration only. It is not financial advice and should not be used as the sole basis for investment decisions.

Table of Contents

Overview

The system predicts the next-day closing price of a stock using historical OHLCV data, S&P 500 market movement, and technical indicators. The default training flow creates one pooled global model across selected tickers and saves it as GLOBAL.pkl. The web client then calls the API with global=true and uses that global model for predictions.

The project also supports training separate per-ticker models such as AAPL.pkl, MSFT.pkl, and TSLA.pkl.

Key Features

  • Next trading day stock closing-price prediction.
  • Historical market-data download with yfinance.
  • S&P 500 daily return as a market-context feature.
  • Technical indicators including moving averages, RSI, Bollinger Bands, volume ratios, spreads, and short-term returns.
  • Default Quantile Gradient Boosting model with prediction ranges.
  • Optional MLP neural-network regressor.
  • Global pooled model across many tickers, with compact ticker identity hash features.
  • Optional per-ticker model training.
  • Flask API with CORS support.
  • Lightweight static HTML/CSS/JavaScript client.
  • Model evaluation using MAE, RMSE, and Pinball Loss for quantile models.

Tech Stack

LayerTechnologies
Machine LearningPython, scikit-learn, NumPy, pandas
Market Datayfinance
Backend APIFlask, Flask-CORS
FrontendHTML, CSS, JavaScript
Model StoragePickle files (.pkl)
Optional TuningOptuna

Project Structure

predictStockMachineLearning-main/
├── Client/
│ ├── CSS/
│ │ └── style.css
│ ├── JS/
│ │ └── script.js
│ └── index.html
├── ModelTraining/
│ ├── features.py
│ ├── model.py
│ ├── predict.py
│ └── train.py
├── Server/
│ ├── requirements.txt
│ └── server.py
├── requirements.txt
├── .gitignore
└── README.md

Main Components

PathPurpose
ModelTraining/features.pyBuilds the feature set used during both training and prediction.
ModelTraining/model.pyContains model wrappers and metric functions.
ModelTraining/train.pyTrains global or per-ticker models and saves them as .pkl files.
ModelTraining/predict.pyLoads trained models and generates next-day predictions.
Server/server.pyExposes the prediction API on localhost:8080.
Client/index.htmlBrowser UI for submitting ticker symbols.
Client/JS/script.jsCalls the Flask API and renders prediction results.

How It Works

  1. Historical stock data is downloaded from Yahoo Finance through yfinance.
  2. S&P 500 historical data is downloaded and converted into daily returns.
  3. Technical indicators are calculated from each ticker's historical price and volume data.
  4. The training script creates supervised examples where today's features are mapped to tomorrow's closing price or tomorrow's return.
  5. A model is trained and saved under ModelTraining/models/.
  6. The Flask server loads the trained model and exposes prediction endpoints.
  7. The web client sends ticker requests to the API and displays current price, predicted price, expected change, and model metrics.

Getting Started

Prerequisites

  • Python 3.10 or newer recommended.
  • Internet connection for downloading market data.
  • A modern browser for the frontend client.

1. Clone the Repository

git clone <your-repository-url>cd predictStockMachineLearning-main

2. Create and Activate a Virtual Environment

On macOS/Linux:

python -m venv .venv
source .venv/bin/activate

On Windows PowerShell:

python -m venv .venv
.venv\Scripts\Activate.ps1

3. Install Dependencies

pip install -r requirements.txt

4. Train a Demo Global Model

This trains one global model on a small demo set: AAPL, MSFT, GOOGL, and TSLA.

python ModelTraining/train.py --demo --target return

The trained model is saved to:

ModelTraining/models/GLOBAL.pkl

5. Start the API Server

python Server/server.py

The API will run at:

http://localhost:8080

6. Open the Web Client

Open this file directly in your browser:

Client/index.html

Enter a ticker symbol such as AAPL, MSFT, GOOGL, or TSLA and click Predict.

Training Models

Train the Default Global Model

By default, if no --tickers or --demo flag is provided, the script attempts to train on the full S&P 500 list.

python ModelTraining/train.py

For a faster demo run:

python ModelTraining/train.py --demo

Train a Global Model on Specific Tickers

python ModelTraining/train.py --tickers AAPL MSFT NVDA AMZN --global-model

Train Separate Per-Ticker Models

python ModelTraining/train.py --tickers AAPL MSFT TSLA --per-ticker

This creates files such as:

ModelTraining/models/AAPL.pkl
ModelTraining/models/MSFT.pkl
ModelTraining/models/TSLA.pkl

Train on Tomorrow's Return Instead of Tomorrow's Price

python ModelTraining/train.py --demo --target return

Training on return can sometimes produce more stable behavior than predicting absolute prices directly. During prediction, the return is converted back into an estimated price.

Use the MLP Neural Network Model

python ModelTraining/train.py --demo --model mlp

Custom hidden layers can be passed as a comma-separated list:

python ModelTraining/train.py --demo --model mlp --mlp-hidden 64,32 --mlp-max-iter 1200

Use Walk-Forward Validation

python ModelTraining/train.py --tickers AAPL MSFT --per-ticker --walk-forward

Tune Gradient Boosting Hyperparameters with Optuna

python ModelTraining/train.py --demo --optuna-trials 25

Optuna is included in the root requirements.txt. Hyperparameter tuning is currently supported for the Quantile Gradient Boosting model.

Running the API Server

Start the server from the project root:

python Server/server.py

The server exposes:

GET http://localhost:8080/health
GET http://localhost:8080/stock?ticker=AAPL&global=true

The server loads models from:

ModelTraining/models/

Using the Web Client

The frontend is a static client located in Client/index.html. It sends requests to:

http://localhost:8080/stock?ticker=<TICKER>&global=true

Because the client uses global=true, make sure ModelTraining/models/GLOBAL.pkl exists before using the UI.

API Reference

Health Check

GET /health

Example response:

{
"status": "ok"
}

Predict Stock Price

GET /stock?ticker=AAPL&global=true

Query parameters:

ParameterRequiredDescription
tickerYesStock ticker symbol, for example AAPL.
globalNoUse the global model when set to true, 1, yes, or y. If omitted, the server attempts to load a per-ticker model.

Example response:

{
"ticker": "AAPL",
"last_close": 195.64,
"last_date": "2026-06-08",
"prediction": 197.21,
"range_low": 192.10,
"range_high": 201.45,
"change": 1.57,
"change_pct": 0.80,
"mae": 3.42,
"rmse": 4.91
}

Response fields:

FieldDescription
tickerNormalized ticker symbol.
last_closeLatest available closing price.
last_dateDate of the latest available market data.
predictionPredicted next-day closing price.
range_lowLower quantile prediction, when available.
range_highUpper quantile prediction, when available.
changeDifference between prediction and latest close.
change_pctPercentage change between prediction and latest close.
maeMean Absolute Error measured during validation.
rmseRoot Mean Squared Error measured during validation.

Modeling Details

Default Model

The default model is a Quantile Gradient Boosting regressor. It trains separate models for multiple quantiles, usually:

0.1, 0.5, 0.9

The median quantile (0.5) is used as the main prediction. The lower and upper quantiles provide an estimated prediction range.

Optional Model

The project also includes a simple MLP regressor based on scikit-learn's MLPRegressor. It uses feature scaling and supports configurable hidden layers.

Features

The default feature set includes:

  • Close
  • SP500_Return
  • SMA_5, SMA_20, SMA_50
  • EMA_5, EMA_20, EMA_50
  • RSI_14
  • BB_Upper_20, BB_Lower_20
  • Volume, Volume_MA_20, Volume_Ratio
  • High_Low_Spread
  • Return_1d, Return_3d, Return_5d

For global models, additional ticker hash features are added by default so that one pooled model can learn ticker-specific patterns without creating one model file per stock.

Metrics

The project reports:

MetricMeaning
MAEAverage absolute prediction error in price units.
RMSESquare-root average squared error; penalizes larger errors more heavily.
Pinball LossQuantile-regression loss used for evaluating quantile predictions.
Baseline MAE/RMSENaive baseline that predicts tomorrow's close as today's close.

Important Notes

  • Generated model files are intentionally excluded from Git by .gitignore.
  • If GLOBAL.pkl does not exist, the web client will not work with the default API request.
  • Training on the full S&P 500 can take significantly longer than the demo mode.
  • Predictions depend on external data from Yahoo Finance, so network issues or unavailable tickers may cause errors.
  • Pickle model files should only be loaded from trusted sources.
  • This is an educational project and not a production trading system.

Future Improvements

  • Add automated tests for feature engineering and API responses.
  • Add Docker support for easier deployment.
  • Add a configuration file for API URL, model type, and default prediction mode.
  • Add charts for historical prices and prediction ranges in the frontend.
  • Add model versioning and experiment tracking.
  • Add CI workflow for linting and test execution.
  • Add a proper LICENSE file before publishing the repository publicly.

License

No license file is currently included in the project. Before publishing or accepting contributions, add a license such as MIT, Apache-2.0, or another license that matches your intended use.

About

100/100 Course advanced programming :)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages