Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Document Ingestion for Search JesterJ logo showing a jester juggling documents, folders, databases and a search bar

A highly flexible, scalable, fault-tolerant document ingestion system designed for search.

LicenseBuild Status

Builds are run on infrastructure kindly donated by

The problem

Frequently, search projects start by feeding a few documents manually to a search engine, often via the "just for testing" built-in processing features of Solr such as SolrCell or post.jar. These features are documented and included to help users get a feel for what they can do with Solr with minimal painful setup.

This is good, and that's how it should be for first explorations. Unfortunately, it's also a potential trap. Large-scale ingestion of documents for search is non-trivial. Many projects outgrow these simple tools and have to throw away their early exploratory work. Nobody likes setting aside valuable work, and it's natural to resist, but the longer one clings to an insufficient tool, the bigger, more difficult, and more expensive the migration is.

Common problems are:

  • It works "ok" for a small test corpus and then becomes unstable on a larger production corpus.
  • The code written to feed into such interfaces (hopefully) reproduces standard solutions to problems that have been solved many times by other search engineers over the last 20 years.
  • No way to recover if indexing errors or is disrupted partway through. One is forced to start again from the beginning.
  • If failure is related to the size of a growing corpus, failures become increasingly common, and eventually, the search index cannot be reindexed or upgraded at all.
  • Leveraging the power of modern multicore machines requires developers skilled at threading and concurrency, the resulting bugs can be very expensive to troubleshoot, fix and test.
  • Reliance on outdated, unmaintained or poorly maintained features such as the Data Import Handler (which has been removed from solr 9.0+). Such features are not used by any major companies (where committers often work), and consequently receive less attention and support.

JesterJ's solution

JesterJ makes it super easy to start with a robust, full-featured indexing infrastructure, so that you don't have to re-invent the wheel, and you don't have to throw away your early work.

The key aspects for achieving this are simplicity, robustness, flexibility, and scalability:

  • A variety of re-usable processing components are provided (flexibility)
  • Scanners (active connectors) for database and filesystem data sources (simplicity)
  • Custom processors only require a 4-method interface (simplicity)
  • Specialized classloading allows any version of a library in your custom code (flexibility, simplicity)
  • Simplified startup: java -jar jesterj.jar <id> <secret>(simplicity)
  • Built in embedded Cassandra for performant persistent storage (simplicity)
  • Optional auto-detection of changes to documents (flexibility, simplicity)
  • Automatic fault-tolerant restart skipping previously seen documents (robustness, scalability)
  • Multithreaded processing to leverage modern machines with large numbers of cores. (scalability)
  • Explicit and direct control of threading. Easy to ensure more threads working on heavy steps (scalability)
  • Single system handling multiple data sources (flexibility, scalability, simplicity)
  • Pre-baked batching of documents for efficient transmission to the search engine (scalability, simplicity)
  • Directed acyclic graph (DAG) capable processing model, and graphical visualization (flexibility, simplicity)

DAG-structured processing is a key feature that is not provided by other tools. Most other tools require a linear pipeline structure, which can become limiting. As time passes, features and enhancements often add complexity. Multiple data sources are also a common dimension for growth. With other systems, you wind up deploying a system per data source. JesterJ is designed to handle complex indexing scenarios.

Consider the following hypothetical indexing workflow, where the system has evolved from a simple linear ingestion into a single index:

  • The source data format changed from, effectively creating a new data source (old data may need reindexing)
  • An external system needed to know that the document was received
  • Product features required a faster, optimized line-item-only search index
  • New features were added to the product that required block-join indexing, but old features couldn't be migrated, so a new index was required.
  • Two new systems also wanted to be notified

In other tools, this will mean six indexing processes (two sources times three indexes), all of which need to send messages, none of which are coordinated if one fails. In JesterJ, it is all one coherent system:

Complex Processing

JesterJ handles such scenarios with a single centralized processing plan, and there is no need to deploy new indexing infrastructure. Furthermore, JesterJ will ensure that if the system is unplugged partway through indexing, you won't get a second message about an order received for everything it processed previously (fault tolerance). The default mode for JesterJ is to ensure at-most-once delivery for steps that are not marked safe or idempotent. Safe steps do not have external effects, and idempotent steps may be repeated en route to the final processing end point.

Getting Started

The best place to start learning more is the documentation in the wiki

Project Status

Current release: 1.0.0 (recommended)

Next Release: 1.1.0

JDK versions

Presently, only JDK 11 has been tested regularly. Unit tests have passed on JDK 17, but the initial system startup and custom class loading are the most JDK-sensitive parts, so we welcome feedback on experiences with more recent JDK versions. Any Distribution of JDK 11 should work. Support for Java 17 and future LTS versions is among our highest priorities for future releases. Building with the latest uno-jar version may be sufficient, but this is not yet certified. nsoft/uno-jar#37

Discord Server

Discuss features, ask questions, etc., on Discord. https://discord.gg/RmdTYvpXr9

Features:

Please see the RELEASE_NOTES.adoc file for full details on features for each available version.

About

Document Ingestion Framework for Search Systems

Topics

Resources

Contributing

Security policy

Stars

37 stars

Watchers

9 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages