Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

WebCrawlAI - AI-Powered Web Scraping Platform

StatusGitHub IssuesGitHub Pull RequestsLicense


AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.


🏆 Sponsors

RapidProxy

RapidProxy Banner
RapidProxy LogoRapidProxy - Stable, scalable residential proxies for efficient data collection.

RapidProxy offers 90M+ real residential IPs across 200+ locations, with high concurrency, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass. Built for developers, data engineers, and businesses working on web scraping, market research, e-commerce, social media automation, and large-scale data operations.

Visit RapidProxy or view the RapidProxy tutorial.

Swiftproxy

Swiftproxy Banner
Swiftproxy LogoSwiftproxy - Seamless access to global web data at scale.

Powered by an 80M+ ethically sourced residential IP pool across 195+ countries, Swiftproxy delivers fast, stable, and secure connections for web scraping, AI, BI, and automation workflows.

Free Trial Available! Experience reliable proxy infrastructure today.

Exclusive: Use code "PROXY90" for 10% off your first purchase!
Visit Swiftproxy

CyberYozh

CyberYozh Banner
CyberYozh LogoCyberYozh - Reliable SMS activation, residential proxies, and mobile proxies for multi-accounting.

We have gathered the best solutions for multi-accounting and automation in one place.

Thordata

Thordata Banner
Thordata LogoThordata - Easy access to web data at scale, perfect for AI.

A global network of 60M+ residential proxies with 99.7% availability, ensuring stable and reliable web data scraping to support AI, BI, and workflows.

🎁 Free Trial Available! Start with our free trial to experience reliable proxy infrastructure.

💰 Exclusive: Use code "THOR66" for 30% off your first purchase!
🔗 Register with invitation code "0HSUJ23G" or click here

📝 Table of Contents

🧐 About

WebCrawlAI is an intelligent web scraping platform designed to help developers, researchers, and businesses extract specific information from websites with ease. The platform combines advanced web scraping capabilities with AI-powered data extraction to handle complex websites, dynamic content, and CAPTCHAs.

The platform features an AI-powered extraction engine that uses Google's Gemini AI model to precisely parse and extract requested information based on natural language prompts. Users can simply provide a URL and describe what data they need (e.g., "Extract all product names and prices") and receive clean, structured JSON output.

Built with modern web technologies, WebCrawlAI emphasizes reliability through robust error handling, retry mechanisms, and comprehensive monitoring. The platform is designed for both technical and non-technical users, providing a user-friendly web interface alongside a powerful API for integration into existing workflows.

🏁 Getting Started

These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.

Prerequisites

  • Python (v3.8 or higher)
  • pip package manager
  • ThorData Residential Proxy account (for web scraping)
  • Google Gemini API Key (for AI-powered extraction)

Installing

  1. Clone the repository

    git clone https://github.com/ArjunCodess/WebCrawlAI.git
    cd WebCrawlAI
  2. Install dependencies

    pip install -r requirements.txt
  3. Set up environment variables Create a .env file and configure the required variables:

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"

    Note: All environment variables are required. The application will raise an error if any ThorData proxy configuration is missing.

  4. Run the application

    python main.py

The application will be available at http://localhost:5000 (default Flask port).

🔧 Running the tests

Currently, the project uses manual testing and user acceptance testing. Automated testing setup is planned for future releases.

Manual Testing

  1. Development Testing

    • Run the development server with python main.py
    • Test core features: web scraping, AI extraction, JSON output
    • Verify error handling and retry mechanisms
  2. Integration Testing

    • Test with various website types (static, dynamic, with CAPTCHAs)
    • Verify AI extraction accuracy with different prompts
    • Test API endpoints and response formats
  3. User Journey Testing

    • Complete web interface workflow
    • Test API integration
    • Verify output format and accuracy

🎈 Usage

Core Features

  1. Web Scraping

    • Handle static and dynamic websites
    • Bypass CAPTCHAs and anti-bot measures
    • Support for JavaScript-heavy sites
  2. AI-Powered Extraction

    • Natural language prompts for data extraction
    • Precise parsing using Gemini AI
    • Structured JSON output
  3. Web Interface

    • User-friendly interface for non-technical users
    • Real-time extraction results
    • Error handling and status updates
  4. API Integration

    • RESTful API for programmatic access
    • Clean JSON responses
    • Easy integration into existing workflows
  5. Monitoring and Analytics

    • Event tracking with GetAnalyzr
    • Performance monitoring
    • Usage analytics

Getting Started Workflow

  1. Access the web interface at the deployed URL
  2. Enter the target website URL
  3. Provide a clear extraction prompt (e.g., "Extract all product names and prices")
  4. Click "Extract Information"
  5. Review the structured JSON output

🚀 Deployment

The project is configured for deployment on Render with the following setup:

Production Deployment

  1. Render Deployment

    • Connect your repository to Render
    • Configure environment variables in Render dashboard
    • Deploy automatically on pushes to main branch
  2. Required Environment Variables

    THORDATA_USERNAME="your_thordata_username"THORDATA_PASSWORD="your_thordata_password"THORDATA_PROXY_SERVER="your_thordata_proxy_server"GEMINI_API_KEY="your_google_gemini_api_key"
  3. Service Configuration

    • Configure as a Web Service on Render
    • Set build command: pip install -r requirements.txt
    • Set start command: python main.py
  4. Monitoring and Error Tracking

    • GetAnalyzr integration for event tracking
    • Built-in error handling and logging
    • Performance monitoring capabilities

Additional Services

  • ThorData Residential Proxy: For reliable web scraping with 60M+ residential proxies
  • Google Gemini AI: For intelligent data extraction and parsing
  • GetAnalyzr: For usage analytics and monitoring

📚 API Documentation

Endpoint:/scrape-and-parse

Method:POST

Request Body (JSON):

{
"url": "https://www.example.com",
"parse_description": "Extract all product names and prices"
}

Response (JSON):

Success:

{
"success": true,
"result": {
"products": [
{ "name": "Product A", "price": "$10" },
{ "name": "Product B", "price": "$20" }
]
}
}

Error:

{
"error": "An error occurred during scraping or parsing"
}

⛏️ Built Using

Core Framework

Web Scraping & Automation

AI & Machine Learning

Frontend & UI

Development & Deployment

Additional Libraries

✍️ Authors

  • ArjunCodess (Arjun Vijay Prakash) - Project development and maintenance

Note: This project embraces open-source values and transparency. We love open source because it keeps us accountable, fosters collaboration, and drives innovation. For collaboration opportunities or questions, please reach out through the appropriate channels.

🎉 Acknowledgements

Sponsors

  • RapidProxy for providing 90M+ residential IPs, smart IP rotation, non-expiring traffic, and AI-powered CAPTCHA bypass for scalable data collection
  • Swiftproxy for providing an 80M+ ethically sourced residential proxy network for web scraping, AI, BI, and automation workflows
  • CyberYozh for providing reliable SMS activation and proxy solutions for multi-accounting and automation
  • Thordata for powering our web scraping infrastructure with their global network of 60M+ residential proxies

Technology Partners

  • Google for providing the Gemini AI model that powers our intelligent extraction capabilities
  • ThorData for reliable residential proxy infrastructure with 60M+ proxies
  • Render for the excellent deployment platform
  • Flask Team for the robust web framework
  • Selenium for powerful browser automation capabilities
  • Open Source Community for the countless libraries and tools that make modern web development possible

WebCrawlAI - Transforming web data into structured insights

Built with ❤️ for developers and data enthusiasts

About

AI-powered web scraping platform that leverages Gemini AI to extract specific information from websites — handles dynamic content, CAPTCHAs, and provides clean JSON output for easy integration.

Topics

Resources

Stars

122 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages