url-normalization
Here are 22 public repositories matching this topic...
Extract and decompose (fuzzy) URLs (including emails, which are conceptually a part of URLs) in texts with Area-Pattern-based modularity
-
Updated
Jan 26, 2025 - TypeScript
Clean, filter and sample URLs to optimize data collection – Python & command-line – Deduplication, spam, content and language filters
-
Updated
Aug 15, 2026 - Python
Canonicalize Varnish cache keys by sorting URL query parameters
-
Updated
Mar 22, 2026 - Roff
Remove clutter from URLs and return a canonicalized version
-
Updated
Jun 3, 2024 - Python
Get a stable, canonical version of any URL, with DNS and HTTPS checks, redirects, tracker stripping, and canonical link extraction!
-
Updated
Aug 14, 2026 - JavaScript
Turn any messy URL into a clean canonical https:// string — or null if invalid. Handles bare domains, www/port stripping, query sorting, punycode, domain extraction, and display-friendly humanization. 980B brotli. TypeScript. ESM + CJS.
-
Updated
Mar 13, 2026 - TypeScript
Allows you to remove ad/tracking query params from a given URL in Scala
-
Updated
Aug 15, 2025 - Scala
🔗 Lightweight Ruby library for parsing URLs and extracting domain components with accurate multi-part TLD support. Handles nested subdomains, query parameters, and URL normalization. Perfect for web scraping, analytics, and URL manipulation. Built on URI and public_suffix gem.
-
Updated
Mar 12, 2026 - Ruby
URL normalizer to canonicalize (standardize) the text representation of a URL to determine if differently-formatted URLs are identical
-
Updated
Mar 26, 2026 - C#
Canonical URL normalization and deterministic SEO payload generation for content platforms.
-
Updated
Jul 20, 2026 - Python
A robust, modular web crawler built in Python for extracting and saving content from websites. This crawler is specifically designed to extract text content from both HTML and PDF files, saving them in a structured format with metadata.
-
Updated
Nov 18, 2024 - Python
Explainable String Intelligence Engine for Splunk. Detect suspicious domains, emails, usernames and DGA-like strings with transparent scoring and explainable detection.
-
Updated
Jul 20, 2026 - Python
🔗 Pathor is a PHP library for normalizing, analyzing, and comparing URLs.
-
Updated
Dec 12, 2024 - PHP
This project contains a Python script to extract all unique absolute URLs from a webpage and write them into a text file. This can be useful for indexing purposes.
-
Updated
Aug 17, 2025 - Python
web crawler from scratch as it possible by using python
-
Updated
Aug 2, 2025 - Python
Natural Language Precessing related notebooks (Machine Learning)
-
Updated
Aug 20, 2021 - Jupyter Notebook
GTM variable template that returns a normalized page path with tracking parameters removed.
-
Updated
May 9, 2026 - Smarty
A simple plugin to rewrite multiple domains or subdomains to a single domain in the site's HTML output.
-
Updated
Jan 22, 2020 - PHP
Improve this page
Add a description, image, and links to the url-normalization topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the url-normalization topic, visit your repo's landing page and select "manage topics."