Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
-
Updated
Aug 15, 2026 - Python
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
Remove clutter from URLs and return a canonicalized version
Canonical URL normalization and deterministic SEO payload generation for content platforms.
A robust, modular web crawler built in Python for extracting and saving content from websites. This crawler is specifically designed to extract text content from both HTML and PDF files, saving them in a structured format with metadata.
Explainable String Intelligence Engine for Splunk. Detect suspicious domains, emails, usernames and DGA-like strings with transparent scoring and explainable detection.
This project contains a Python script to extract all unique absolute URLs from a webpage and write them into a text file. This can be useful for indexing purposes.
web crawler from scratch as it possible by using python
Synthetic-data pipeline for URL normalization, source extraction, checkpointing, and GEO citation analysis.
Add a description, image, and links to the url-normalization topic page so that developers can more easily learn about it.
To associate your repository with the url-normalization topic, visit your repo's landing page and select "manage topics."