Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

274 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

scraper

Project for scraping content out of pages and/or feeds.

The big idea here is to use fluent builder to make a simple scraping DSL. For example, the simplest scraping job would look like this:

String pageContent = new Scraper.Builder().url("http://www.apple.com").getResult();

Chaining calls to do things like stipulate whether you want just text, or if you want any transforms performed, so the above could be changed like this to return just the text (for instance for a classification engine):

String pageContent = new Scraper.Builder().url("http://www.apple.com").asText().getResult();

Further manipulators could be used to do things like direct the scraper to certain elements, e.g. suppose we wanted the 3rd table as HTML and nothing more, something like this:

List<String> urls = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().getResults())
.getResults();

Here is a more advanced case. We want to get the 3rd table, extract the links from it, then get the value of the parameter oppId from each link:

List<String> ids = scraper
.url(testTableHtmlUrl)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();

Note that the keys are collected in the getResults() method in the extractor, but another getResults() call is needed in the scraper because we might have to iterate, in which case each page would have an extraction and the results would be collected.

Iteration

In a lot of cases, you want to scrape something from a page, but the same form is repeated on multiple pages. To support that, we have the notion of an iterator. The scraper will call the iterator each time it's ready for a new page. All the iterator has to do is construct the URL for the next page. Like this:

	Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
List<String> ids = scraper
.url(testTableHtmlUrl)
.pages(1)
.iterator(pageIterator)
.extract(scraper.extractor().table(3).links().parameter("oppId").getResults())
.getResults();
assertThat(ids.size(), is(86));

Notice that we are constraining the iteration with the pages method. We probably want to support an open-ended iteration where the scraper will keep trying to get more pages until it gets a 404 and then it will exit. This is necessary because we may not know how many pages there are and pages may be added at some point. (Implementing this is not very difficult: inside the scraper, it sets up the extractor, gets the results, then checks if there is an iterator and if there is, it calls it in a loop, collecting all the results.)

Listing and Detail: Following Links to a Detail Page

Another common scenario is that you have a set of links that you have to follow to a detail page where the actual content is that you want to scrape. That's what this syntax is meant to support. Here is an example:

@Test
public void useIteratedListingAndDetailInterface() throws IOException {
Scraper scraper = new Scraper();
Iterator pageIterator = new Iterator() {
@Override
public URL build(int i) {
String nextPageUrl = MessageFormat.format("/testpages/ids-page-{0}.html", i + 2);
log.debug("next page to iterate to: {}", nextPageUrl);
return TestUtil.getFileAsURL(nextPageUrl);
}
};
Scraper detailScraper = new Scraper();
List<Map<String, String>> records = scraper
.url(testTableHtmlUrl)
.pages(3)
.iterator(pageIterator)
.listing(scraper.extractor().table(3).links().getResults())
.detail(detailScraper)
.getRecords();
assertThat(records.size(), is(greaterThan(0)));
log.debug("fields = {}", records);
}

Notice that we have to have a separate scraper for extracting the details.

About

For scraping content out of pages and/or feeds.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors