Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

MRoS

MapReduce on Steroids - Project for the Distributed Systems lecture at DHBW Karlsruhe

Spoiler: This implementation of MapReduce isn't actually on steroids, as for the following reasons:

  1. We don't support the use of drugs of any kind.
  2. The goal of this project is to re-implement the fundamental concept of how MapReduce works in an easy to understand manner. Actually going for performance would require very complex algorithmic implementations as well as the use of a distributed file system.
  3. The developers of known implementions such as Apache Hadoop are way more knowledgable of the topic and "might" have got a few more resources and time at their disposal.

Still, this project is a good visualization of the fundamental principle of MapReduce. Take a look into the code - I promise it's easy to understand. 😉

Requirements

  • Python (tested on 3.10)
    • Packages: dill for serialization

Limitations

Our prototype is still in the early stages and has several limitations, but we tried to stay true to the original concept where possible and viable given the short time we had. The algorithmic idea of MapReduce is simple, but implementing it properly takes careful planning and requires optimizations.

Our current prototype is just a rough outline and doesn't have all the features yet. One limitation is that it does not rely on distributed filesystems, which the original does. Hence, there is no network of intertwined nodes that can self-orchestrate. Instead, we rely on a setup where a controller starts and manages several worker processes. The controller is responsible for accepting client requests. Those contain the data along with the map and reduce functions they want to see applied. The controller then splits the data into fair chunks for each worker, sends them over, gathers the results (in multiple steps) and transmits it back to the requesting client.

In addition, establishing communication between the workers is not a straightforward task. In the original implementation, the mapper processes would distribute the resulting data to the reducers based on their corresponding keys. However, we have simplified this process in our prototype. Instead, the controller sends the data and map function in a first step, then the workers perform the mapping in parallel. The result is sent back to the controller, which then shuffles the data and redistributes it to the reducers. Using this approach, no interaction between the workers must be implemented. However, it also means that multiple transmissions to the controller are needed.

The system currently allows for only one pass of map and reduce for each request to the controller. More complex scenarios can be realized by submitting multiple "independent" requests to the controller, each with its own map and reduce function. An example of which is shown in the matrix_multiplication.py script. Furthermore, if there is a requirement for more complex scenarios that involve multiple map or reduce steps, our system accommodates this as well. In such cases, the client can access the controller multiple times, each time with different map and reduce functions specified. An example of this can be seen in the matrix_multiplication.py script.

Despite its limitations, the implementation is capable of effectively processing data sets by leveraging the logical separation of map and reduce functions. To ensure fairness, the data is divided into manageable chunks, ensuring that each worker receives an equitable workload. The processing follows a sequential pattern of map operations followed by reduce operations. Although the implementation is rudimentary, it allows for basic data processing using the MapReduce paradigm.

Usage

  1. Install requirements: pip install -r requirements.txt
  2. Start the master (e.g. with five workers - ports are assigned automatically):
    1. python master.py --host localhost --port 8000 --worker-host localhost --worker-amount 3

This configuration describes a master (orchestrator) that listens to localhost:8000. It manages 3 workers that are also hosted on localhost, their ports are automatically assigned.
NOTE: Ensure that you start the master using the plain python command, as the workers are started using that as well. If you start the master using e.g. python3, the correct working of the script cannot be ensured.

  1. Describe your MapReduce request, such as shown in the word counter example: python word_counter.py
  2. Enjoy, take a sip of your favorite coffee and wait for your results (if you actually spent time generating a sufficiently large example that you have to wait for an answer)!

Examples

We've included two examples on how a MapReduce-request could be sent using our MapReduce prototype:

The word counter is one of the most trivial examples for show-casing MapReduce. The matrix multiplication example on the other hand, is a bit more sophisticated as it requires two separate MapReduce-requests - one for retrieving all required cell products and the other for summing those products up for each cell in the final matrix.

About

MapReduce on Steroids - Project for Distributed Systems lecture at DHBW Karlsruhe

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages