Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

OpenPrice AI Scraper

OpenPrice AI Scraper

A company-neutral, open-source desktop and command-line application for extracting candidate product prices from public web pages and exporting an auditable Excel report.

The repository contains no customer spreadsheet, supplier list, operational log, private technical manual, browser installer, web-driver binary, API key, proxy credential, or customer branding. It starts from a blank input template and is designed for reuse by any organization that is authorized to collect the target information.

Responsible-use notice: only collect pages that you are authorized to access. Comply with applicable law, website terms, robots.txt, rate limits, privacy requirements, and contractual restrictions. This software does not include features to bypass authentication, CAPTCHAs, paywalls, or technical access controls.

What it does

  • imports neutral .xlsx, .xlsm, or .csv product lists;
  • accepts direct product URLs or performs optional public URL discovery;
  • uses requests first and optional Selenium/Chrome fallback for rendered pages;
  • supports an optional authenticated HTTP proxy, including Bright Data credentials;
  • extracts candidates from JSON-LD, product metadata, and visible DOM elements;
  • validates candidates with a local model, OpenAI, Anthropic, or Firecrawl;
  • keeps secrets in the operating-system keychain or environment variables;
  • flags extreme candidate prices with a robust MAD rule;
  • exports URL-level evidence, selected prices, input snapshots, summary metrics, and safe run metadata;
  • includes a local demonstration site, tests, a public-repository audit, and an MIT license.

OpenPrice architecture

Repository structure

ProcureLens-AI/
├── assets/ # Neutral logo and architecture diagram
├── config/ # Empty allowlist/denylist examples
├── docs/ # Architecture, mathematics, providers, privacy, and legal use
├── examples/ # Offline-safe end-to-end demonstration
├── scripts/ # Environment check, demo server, reset, and audit tools
├── src/openprice_scraper/ # Application source code
│ └── providers/ # Local, OpenAI, Anthropic, and Firecrawl integrations
├── templates/ # Blank XLSX and CSV input templates
├── tests/ # Unit tests and synthetic HTML fixture
├── LICENSE # MIT License
├── pyproject.toml # Package and optional dependency groups
└── README.md

Requirements

Minimum

  • Python 3.11 or newer;
  • internet access for live scraping;
  • permission to access the selected pages.

Optional components

  • Chrome or Chromium for Selenium fallback;
  • an OS keychain backend for encrypted-at-rest credential storage;
  • an OpenAI, Anthropic, or Firecrawl account for the corresponding provider;
  • a Bright Data or other HTTP proxy account when proxy routing is required;
  • scikit-learn for training the optional local Random Forest;
  • PySide6 for the desktop interface.

Selenium 4 includes Selenium Manager, which automatically resolves a compatible driver when no driver is supplied. The repository intentionally does not bundle ChromeDriver, GeckoDriver, browsers, installers, or executables.

Installation

macOS

xcode-select --install 2>/dev/null ||true
brew install python@3.12
cd~/Downloads/ProcureLens-AI
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Install Chrome only when browser fallback is required:

brew install --cask google-chrome

Windows PowerShell

Install Python 3.11+ from Python.org or winget, then:

cd $HOME\Downloads\ProcureLens-AI
py -3.12-m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[full]"

Install Chrome only when browser fallback is required:

winget install Google.Chrome

Ubuntu / Debian

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
cd~/Downloads/ProcureLens-AI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[full]'

Optional Chrome/Chromium package names vary by distribution. Selenium Manager can manage the driver, but a browser must still be available.

Minimal command-line installation

This omits the desktop UI and cloud/ML providers:

python -m pip install -e '.[secure,browser]'

Feature-specific installation

python -m pip install -e '.[gui]'# desktop interface
python -m pip install -e '.[openai]'# OpenAI provider
python -m pip install -e '.[anthropic]'# Anthropic provider
python -m pip install -e '.[firecrawl]'# Firecrawl rendering
python -m pip install -e '.[ml]'# Random Forest training

Check the environment:

python scripts/check_environment.py

First-run configuration

Run the setup wizard once:

openprice setup

It configures:

  • local or cloud candidate provider;
  • default market and currency;
  • number of parallel workers and URLs per item;
  • per-domain delay and browser fallback;
  • robots.txt behavior;
  • optional proxy host, port, username, and password;
  • output directory.

API keys and proxy passwords are stored with keyring in the operating-system keychain. They are not written to config.json, workbooks, logs, or source files.

Environment variables can be used instead:

cp .env.example .env
# Load variables with your preferred shell or secret manager.

Recognized secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
FIRECRAWL_API_KEY
BRIGHTDATA_USERNAME
BRIGHTDATA_PASSWORD
OPENPRICE_PROXY_URL

Never commit a populated .env file.

Prepare the product file

Use either:

  • templates/parts_template.xlsx; or
  • templates/parts_template.csv.

The workbook is intentionally blank. The main sheet is named Items.

ColumnRequiredExample meaning
item_idRecommendedYour neutral internal row identifier
supplier_referenceRecommendedManufacturer or supplier part number
manufacturerOptionalBrand/manufacturer
descriptionRecommendedProduct description
countryOptionalTwo-letter market code such as US
currencyOptionalISO code such as USD
target_urlOptionalOne or more semicolon-separated product URLs
enabledOptionaltrue or false
notesOptionalNon-secret note

At least one of description, supplier_reference, or target_url must be populated. Direct product URLs are more reproducible and create less search traffic.

Validate before scraping:

openprice inspect templates/parts_template.xlsx

More details: Input format.

Run from the command line

openprice scrape /path/to/parts.xlsx

Specify the output and provider:

openprice scrape parts.xlsx \
--output output/prices.xlsx \
--provider local \
--workers 3 \
--max-urls 5

Cloud provider example after openprice setup:

openprice scrape parts.xlsx --provider openai

Start the desktop application

openprice gui

The Run tab imports and previews the product file. The Settings tab stores provider keys and non-secret settings. The report location is selected before the run begins.

Offline end-to-end demonstration

No external website or API key is needed.

Terminal 1:

source .venv/bin/activate
python scripts/run_demo_server.py

Terminal 2:

source .venv/bin/activate
openprice scrape examples/demo_items.csv \
--provider local \
--output output/demo-results.xlsx

The local page contains a synthetic product and price. Open output/demo-results.xlsx to inspect the selected price and extraction evidence.

Output workbook

The generated workbook contains:

  • Results — one row per evaluated URL;
  • Best Prices — the highest-confidence non-outlier candidate for each item;
  • Input Snapshot — canonical copy of input rows;
  • Summary — item counts, observations, selected prices, and outlier count;
  • Run Metadata — safe settings only, with secret-like keys removed.

The report is evidence for review, not a contractual quotation. Taxes, shipping, minimum order quantities, regional availability, and negotiated pricing may not be visible on the page.

Local AI model

The default provider stays on the user's machine. It creates a feature vector for each amount and computes a logistic confidence:

$$ z_i=b+\mathbf{w}^{\mathsf T}\mathbf{x}_i, \qquad P_i=\frac{1}{1+e^{-z_i}}. $$

Features include structured-data source, currency agreement, price/list-price attributes, product-token overlap, sale terms, logarithmic price magnitude, and context length.

An optional Random Forest replaces the heuristic after training:

$$ \widehat P(y=1\mid\mathbf{x})= \frac{1}{T}\sum_{t=1}^{T}P_t(y=1\mid\mathbf{x}). $$

Train with your own legally collected labels:

openprice train candidate_labels.jsonl

See Mathematical model and Local model training.

Robust price cleaning

For several observations of the same item and currency, the report calculates the median m and median absolute deviation:

$$ \text{MAD}=\text{median}(|p_i-m|). $$

A candidate is flagged when the modified z-score exceeds 3.5:

$$ |z_i^*|= \left|0.6745\frac{p_i-m}{\text{MAD}}\right|>3.5. $$

This is a review aid. It can incorrectly reject valid prices in small or multimodal markets.

Providers

ProviderPage fetching / validationPage text leaves the computer?API key
localRequests/Selenium + local rankingNoNo
openaiLocal extraction + OpenAI candidate validationYes, bounded excerptYes
anthropicLocal extraction + Claude candidate validationYes, bounded excerptYes
firecrawlFirecrawl rendering + local rankingYes, target URLOptional/current plan dependent

Read Provider details before enabling cloud processing.

Proxy and Bright Data

Proxy routing is optional. Enter the exact host, port, username, and password issued by the provider. Credentials are stored in the OS keychain. The requests fetcher supports authenticated proxy URLs.

Authenticated Chrome proxies require a local gateway or provider-specific browser integration. The application does not expose proxy credentials in browser command-line arguments.

See Proxy and Bright Data configuration.

Testing

python -m unittest discover -s tests -v
python scripts/audit_public_repo.py

With developer tools:

python -m pip install -e '.[dev]'
ruff check .

The tests use synthetic pages and temporary workbooks. They do not access the public internet.

Security and sanitization guarantees

The public repository is designed to reject common accidental disclosures:

  • no hard-coded API keys or fallback credentials;
  • no absolute personal user paths;
  • no input/output workbooks except the blank template;
  • no logs, IP histories, cookies, supplier/domain datasets, or scraped HTML archives;
  • no customer logo, icon, trademark, or internal documentation;
  • no browser installer, driver, proxy-manager installer, .exe, .dll, or .lnk;
  • no serialized private model or training dataset;
  • secret-like metadata is excluded from output reports;
  • scripts/audit_public_repo.py checks common secret and binary patterns in CI.

If any credential was previously stored in another source tree or archive, revoke and rotate it before publishing that history.

Reset local data

python scripts/reset_local_data.py --yes

This deletes local settings, model files, generated outputs, and keychain entries associated with OpenPrice. It does not delete repository files or unrelated browser data.

Known limitations

  • Website layouts and structured data change over time.
  • Search discovery can miss the correct product or return an accessory.
  • Browser fallback is slower and more resource-intensive than direct HTTP.
  • Cloud models can make errors and incur costs.
  • Currency is detected and matched against the requested currency, but amounts are not converted between currencies (e.g. EUR to USD).
  • A public page price may be stale, personalized, excluding taxes, or conditional on quantity.
  • The default local model uses fixed, hand-tuned weights and performs well in practice; training the optional Random Forest on your own labeled data gives a formally validated, domain-specific alternative.

Documentation

License

Released under the MIT License.

Author

Dev Kumar.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages