Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - jaron/sciencegraph: A comprehensive knowledge graph of scientific concepts · GitHub
Skip to content

Latest commit

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ScienceGraph

A comprehensive knowledge graph of scientific concepts, for question answering

Introduction

The ScienceGraph knowledge graph is distributed as a Neo4j database, and currently consists of about 1.4 million nodes and 2.4 million relationships. Much of the data comes from the general knowledge ConceptNet 5 project, which has been curated to filter out much of its non-scientific content, on order to produce a more focused, domain-specific body of knowledge that will provide a more accurate basis for question answering.

ScienceGraph was compiled for the purpose of answering scientific questions using natural language. This repo includes two collections of sample questions graciously released by the Allen Institute of AI, containing science exam questions posed to 11 and 15 year-olds. Instructions for running the question answering code are provided below.

Getting Started
  • If you haven't already, download and install the Neo4j graph database (The free community edition is fine)
  • Begin by cloning this repo, cd data and uncompress the database contents with sh extract.sh - this should result in a directory called sciencegraph.db
  • Now start the Neo4j server, and when prompted for the Database Location, select the sciencegraph.db directory you have just unzipped. Then click Start, and wait for a few moments whilst the server initialises.
  • The Neo4j server dialog will report when the server is ready, you can know access your graph database's web interface at http://localhost:7474/browser/
  • In the web interface, try running a sample Cypher query like start n=node(*) match n return count(n) - if everything is working you should see a number that's greater than 1300000
  • Now, build the java project using mvn package
  • You can now try using the database to answer a single question chosen at random using java -cp target/sciencegraph-1.0-jar-with-dependencies.jar com.agentsmith.sciencegraph.App -once
  • Use the same command without the -once parameter to create answers for all the questions in the set, this might take about 5-10 minutes
How It Works

The basic stages of the current question answering process are as follows:

  • The plain text of a question is passed to the ProblemSolver.solve() method
  • A Blackboard instance is created, this will store intermediate results produced by the various stages of the comprehension pipeline
  • GraphQueryService.identifyQuestionConcepts() identifies concepts in the question that also have corresponding nodes in the knowledge graph, this filters out irrelevant words and concepts
  • GraphQueryService.identifyAnswerConcepts() is called next and does the same with each of the multiple choice answers
  • QuestionClassifier.classify() uses rules to determine the type of the question
  • Knowing the question type then allows the answering process to select a tailored strategy for that particular type
  • The default answering strategy is to determine the semantic closeness between each concept in the question and each concept in the answer, and compute the average shortest path between them.
  • The scoring system then evaluates the computed path lengths, to find the answer which is conceptually closest to the question
  • There is also some primitive weighting applied before the final answer is selected. This an area that I aim to improve by incorporating some machine learning
When the answer is wrong

"The most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' but 'That's funny...'" -- Isaac Asimov

The question answering logic is still a work in progress, so don't be surprised if it produces the wrong answer. But like Asimov said, wrong answers can provide the beginnings of great insights. So you might like to look at the console output, and see if you can identify why the system got it wrong.


Acknowledgements

ScienceGraph includes data from ConceptNet 5, which was compiled by the Commonsense Computing Initiative. ConceptNet 5 is freely available under the Creative Commons Attribution-ShareAlike license (CC BY SA 3.0).

About

A comprehensive knowledge graph of scientific concepts

Topics

Resources

Stars

16 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages