Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

History
204 lines (147 loc) · 8.12 KB

File metadata and controls

204 lines (147 loc) · 8.12 KB

python

Some small Python utilities.

To use them, make sure you also downloaded util.py file from the same directory (or better, just clone the whole repo).

instalive.py

Instagram live downloader. Feed in MPD URL (including all the query parameters!) within 24 (a few) hours and it can download the full live stream from the beginning.

Usage:

usage: instalive.py [-h] [--action ACTION] [--dir DIR] [--debug] [--quality QUALITY] [--time TIME] [--range RANGE] url
Available actions:
all - Download both video and audio (including live and backtracking), and then merge them (default)
live - Download the live stream only (no backtracking)
video - Download video only
audio - Download audio only
merge - Merge downloaded video and audio
check - Check the downloaded segments to make sure there are no missing segments
manual - Manually check missing segments at the largest gap (or use --range to assign a range) (for debugging only)
info - Display downloader object info
import:<path> - Import segments downloaded via N_m3u8DL-RE from a given path
positional arguments:
url url of mpd
options:
-h, --help show this help message and exit
--action ACTION, -a ACTION
action to perform (default: all)
--dir DIR, -d DIR save path (default: instalive_{mpd_id})
--debug debug mode
--quality QUALITY, -q QUALITY
manually assign video quality by quality name.
(default: "pst": which is the "original" (?) and has the best bitrate,
but not necessarily the highest resolution.
Pass empty string or "highest" to use the highest resolution one.)
--time TIME, -t TIME for debugging only; manually assign last t (default: auto)
--range RANGE for debugging only; manually assign iteration range (start,end) for manual action

oricon.py

Deprecated: You can no longer get original size of images from Oricon website AFAIK. The random quality toggling because of CDN is also gone (?). This can still be used to download the images in the size shown on the webpage, though.

Quickly download all the highest quality pictures from any Oricon news, special or photo article.

Usage:

oricon.py https://www.oricon.co.jp/news/2236438/

scraper_ameblo_api.py

Ameblo (アメーバブログ or アメブロ, Japanese blog service) downloader. Supports images and text.

Usage:

CLI:

usage: scraper_ameblo_api.py [-h] [--theme THEME] [--output OUTPUT] [--until UNTIL] [--type TYPE] blog_id
Download ameblo images and texts.
positional arguments:
blog_id ameblo blog id
optional arguments:
-h, --help show this help message and exit
--theme THEME ameblo theme name
--output OUTPUT, -o OUTPUT
folder to save images and texts (default: CWD/{blog_id})
--until UNTIL download until this entry id (non-inclusive)
--type TYPE download type (image, text, all)

As Python module:

fromscraper_ameblo_apiimportdownload_alldownload_all('user_id', save_folder='.', limit=100, download_type='all')

scraper_fantia.py

Fantia downloader. Inspired by dumpia.

Usage:

fromscraper_fantiaimportFantiaDownloaderkey='your {_session_id}'# copy it from cookie `_session_id` on fantia.jpid=11111# FC id copied from URLdownloader=FantiaDownloader(fanclub=id, output=".", key=key)
downloader.downloadAll()

Or just download certain post (you can omit fanclub id in this case):

downloader=FantiaDownloader(output=".", key='your _session_id')
downloader.getPostPhotos(12345)

scraper_radiko.py

Deprecated: just use yt-dlp or even better, with yt-dlp-rajiko plugin.

Note: you need to prepare the JP proxy yourself.

Usage:

fromscraper_radikoimportRadikoExtractor# It supports the following formats:# http://www.joqr.co.jp/timefree/mss.php# http://radiko.jp/share/?sid=QRR&t=20200822260000# http://radiko.jp/#!/ts/QRR/20200823020000url='http://www.joqr.co.jp/timefree/mss.php'e=RadikoExtractor(url, save_dir='/output')
e.parse()

nico.py

Nico Timeshift downloader. Download both video and comments. Also can download thumbnail from normal video. Downloading for normal video isn't supported (yet).

pip install browser-cookie3 websocket-client rich python-dateutil pytz requests beautifulsoup4 lxml

The actual downloading is delegated to minyami and/or N_m3u8DL-RE, so make sure you have them installed first and available in your PATH.

It also can reads the arguments from a nico.txt file in the CWD or the same directory as the script. Syntax: just put all the arguments in one line.

CLI:

usage: nico.py [-h] [--verbose] [--info] [--dump] [--thumb] [--cookies COOKIES] [--comments {yes,no,only}] [--proxy PROXY] [--save-dir SAVE_DIR] [--reserve]
[--simulate]
url
positional arguments:
url URL or ID of nicovideo webpage
options:
-h, --help show this help message and exit
--verbose, -v Print verbose info for debugging.
--info, -i Print info only.
--dump Dump all the metadata to json files.
--thumb Download thumbnail only. Only works for video type (not live type).
--cookies COOKIES, -c COOKIES
Cookie source.
Provide either:
- A browser name to fetch from;
- The value of "user_session";
- A Netscape-style cookie file.
--comments {yes,no,only}, -d {yes,no,only}
Control if comments (danmaku) are downloaded. [Default: no]
--proxy PROXY Specify a proxy, "none", or "auto" (automatically detects system proxy settings). [Default: auto]
--save-dir SAVE_DIR, -o SAVE_DIR
Specify the directory to save the downloaded files. [Default: current directory]
--reserve Automatically reserve timeshift ticket if not reserved yet. [Default: no]
--simulate Simulate the download process without actually downloading.

util.py

Some utility functions mainly for myself. Read the code to get the idea.

Some highlights:

Network & Web Scraping:

  • download - a comprehensive file downloader with retry logic, duplicate handling (skip/overwrite/rename), referer support, and automatic filename detection from URLs or response headers
  • get - a convenience wrapper around requests that returns a BeautifulSoup object with retry logic built-in
  • requests_retry_session - create a requests session with automatic retry on network failures
  • load_cookie - load cookies from browser (Chrome/Firefox/Edge), cookie files (Netscape format), or cookie strings

File Operations:

  • safeify - sanitize filenames by replacing illegal characters with full-width equivalents (yes I know safeify isn't a real word. It was blindly copied from another project and It was too much effort to change it)
  • move_or_delete_duplicate - move files with smart duplicate detection via hash comparison
  • dump_json/load_json - convenient JSON file I/O with proper encoding and formatting

Data Structures & Formatting:

  • Table - a simple table class with sorting, searching, filtering, and pretty printing capabilities
  • format_str - format a string to certain width and alignment. Supports wide characters like Chinese and Japanese
  • compare_obj - recursively compare two objects (dicts, lists, etc.) and print differences with rich formatting
  • rprint - print function using the rich library for colored and styled terminal output

Date & Time:

  • MyTime - a datetime wrapper that makes timezone conversions easy (local, JST, naive)
  • parse_to_shortdate - parse a date string to a short date string ('%y%m%d'). Supports more East Asian language formats than dateutil.parser