This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
This repository was archived by the owner on Jul 15, 2019. It is now read-only.

Latest commit

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Cognitive Client

DEPRECATED: this repo is no longer actively maintained. It can still be used as reference, but may contain outdated or unpatched code.

This project contains several methods to make it easy to use Watson Developer Cloud services, particularly NaturalLanguageUnderstanding and AlchemyLanguage, to analyze multiple documents. Some of the key features include:

  • Performing Web searches and feeding the results directly to the Watson Developer Cloud. The user can specify the number of documents to search for, as well as whether to search the entire Web or only news stories.
  • Aggregating results across multiple documents. These documents might have been obtained via Web searches, although this does not have to be the case.
  • Storing documents from Web searches locally so they can later be accessed quickly. These locally stored documents can easily be analyzed by the Watson Developer Cloud.
  • Allowing all files in a directory to be easily analyzed using the Watson Developer Cloud with the results aggregated.
  • Allowing analyzed data to be easily stored and retrieved from disk, as well as combined.
  • Allowing directories of data analysis files to easily be read, written, and aggregated.
  • Convenient methods to obtain data analysis values and statistics.
  • Methods for aggregating quantities such as sentiment analysis values across several documents.

Getting Started

Installation

Maven

If Apache Maven is being used, the following dependency should be included:

 <dependency>
<groupId>com.ibm.watson.developer_cloud</groupId>
<artifactId>cognitive-client-java</artifactId>
<version>1.0</version>
</dependency> 
Gradle

If Gradle is being used, the following dependency should be included:

 compile 'com.ibm.watson.developer_cloud:cognitive-client-java:1.0'

Using the Cognitive Client

The following classes should be imported as needed:

importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData;
importcom.ibm.watson.developer_cloud.cognitive_client.AggregateData.Data;
importcom.ibm.watson.developer_cloud.cognitive_client.AlchemyClient;
importcom.ibm.watson.developer_cloud.cognitive_client.DataManager;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageClient;
importcom.ibm.watson.developer_cloud.cognitive_client.NaturalLanguageUnderstandingClient;
importcom.ibm.watson.developer_cloud.cognitive_client.Search.SearchType;
importcom.ibm.watson.developer_cloud.cognitive_client.Util;
importcom.ibm.watson.developer_cloud.cognitive_client.Util.DataType;

There are two natural language services that can used with our cogitive client: Alchemy and Natural Language Understanding which are both available from IBM's Watson Developer Cloud. The following creates a client for Alchemy:

NaturalLanguageClientclient = newAlchemyClient(apikey);

The following creates a client for NaturalLanguageUnderstanding:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password);

In some cases, it is desirable to limit the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding. The following creates a client which limits the number of entities, keywords, and concepts returned by a single call to NaturalLanguageUnderstanding to 5:

NaturalLanguageUnderstandingClientclient = newNaturalLanguageUnderstandingClient(userid, password, 5);

The variable "client" defined using the constructors above can be used to analyze data using the Watson Developer Cloud as illustrated below.

The following calls the Watson Developer Cloud to get combined analysis including concepts, entities, keywords, and categories/taxonomies:

AggregateDataad = client.analyzeData("https://en.wikipedia.org/wiki/IBM", DataType.URL, "IBM Wikipedia entry");

The 2nd parameter indicates whether the 1st parameter is a url, text data, or html data. The 3rd parameter is a string provided by the user which gives an explanation of the data set.

The following:

AggregateDataad = DataManager.analyzeWebSearchResults ("IBM", 50, SearchType.GOOGLE_REGULAR, "IBM Google search", false, false, null, client);

performs a search on "IBM". The top 50 search results are analyzed by the Watson Developer Cloud with the results being stored in "ad". The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The 4th parameter is a string provided by the user which gives an explanation of the data set. The 5th parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "false" indicates that "ad" will only contain a summary of the results, and not the analysis results from each individual document. The 6th parameter indicates whether or not the analysis results for each document should be stored on disk. If the 6th parameter is true, the 7th indicates the directory for storing the analysis results.

"ad" is serializable, with an implemented toString method. In order to see the contents of ad, use

System.out.println(ad);

The following returns an ArrayList of the most frequently occurring disambiguated entities sorted by decreasing frequency of occurrence:

ArrayList<Entry<String,Data>> sortedCounts = ad.getSortedValues(AggregateData.Type.DISAMBIGUATEDENTITY,AggregateData.DataType.COUNT);

The following returns an ArrayList of the most relevant keywords occurring sorted by decreasing sum of relevancy scores:

ArrayList<Entry<String,Data>> sortedRelevancy = ad.getSortedValues(AggregateData.Type.KEYWORD,AggregateData.DataType.RELEVANCE);

The following writes "ad" to "filename" as binary data:

ad.writeToFile(filename);

The following reads in "ad2" from the binary file "filename":

AggregateDataad2 = AggregateData.readFromFile(filename);

It is also possible to analyze an entire directory of files. The following analyzes all files in "dir3" (but does not recursively search subdirectories):

AggregateDataad = DataManager.analyzeDirectory("dir3", "IBM search results", true, true, "dir3-analysis", DataType.HTML, client);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether "ad" should contain results from the analysis of all analyzed documents, or only the summary results. The fact that it is "true" indicates that "ad" will contain the text data analyzed as well as the analysis results from each individual document, in addition to the summary results. The 4th parameter indicates whether or not the analysis results for each document should be stored on disk. Since the 4th parameter is true, the 5th parameter indicates the directory for storing the analysis results. The 6th parameter indicates whether the files being analyzed are text, html, or each contain a url representing a Web document to be analyzed.

Supposing the directory "dir3-analysis" contains files from analyzing text documents. The following method call aggregates all of the result files contained in this directory:

AggregateDataad = DataManager.aggregateDirectoryStats("dir3-analysis", "IBM search results", false);

The 2nd parameter is a string provided by the user which gives an explanation of the data set. The 3rd parameter indicates whether analysis results read in from each file should be stored in the returned data structure. The fact that it is false means that "ad" will only contain a summary of the results, and not the actual results stored in each file.

In some cases, it is desirable to add the results from one AggregateData structure to another:

data.combineData(newdata, true); // "data" and "newdata" are both of type "AggregateData"

This method call combines the data stored in "newdata" with the data stored in "data". The 2nd parameter indicates whether or not raw data stored in "newdata" should be added to raw data stored in "data". Since it is true, raw data stored in "newdata" is added to raw data stored in "data".

In some cases, it is desirable to perform a search and store all of the documents returned from the search on disk. That way, it is not necessary to re-fetch the documents from the Web when they need to be viewed more than once. In addition, the documents stored on disk can subsequently be passed to the Watson Developer Cloud to analyze their contents. This can be achieved via the following:

Util.searchWeb("IBM", 15, SearchType.GOOGLE_REGULAR, "dir2", ".html");

In this example, a search for the first 15 responses to the query "IBM" is performed. The 3rd parameter indicates the type of search. SearchType.GOOGLE_REGULAR indicates a regular Google search. SearchType.GOOGLE_NEWS would indicate a Google search of just news stories. The files are stored in the directory "dir2". The last parameter is a suffix assigned to the file names. Each file name is the urlencoded version of the url appended with the last parameter.

About

DEPRECATED: this repo is no longer actively maintained

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages