Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

magi - cloud storage MAnaGement kIt

This software provides several functions for managing large-scale cloud object store, leveraging Apache Spark. It uses AWS Java SDK and Azure Storage Java SDK to communicate with Amazon AWS, Microsoft Azure or NetApp StorageGrid.

One specific use case for NetApp StorageGrid customers is to use this software to replicate or migrate objects between a StorageGrid system and Microsoft Azure or Amazon S3.

Functions Implemented

  • Bucket synchronization
    Replicate or migrate objects between two buckets in AWS, Azure or NetApp StorageGrid
  • Empty a bucket
    Delete all objects in a bucket
  • List a bucket
    List all objects in a bucket
  • Fill a bucket
    Populate a bucket with objects

Software

How to compile from source code?

  1. Compile magi
$ git clone https://github.com/NTAP/magi.git $ cd magi
$ sbt package
$ ls target/scala-2.11/magi_2.11-1.2.jar

The jar file is ready.

How to run it?

  1. You need a spark cluster up and running, either at Microsoft Azure, AWS EC2 or your own data center.
  2. Install the two Java SDK jars in your spark cluster (assume spark is installed at /usr/spark.).
    $ scp aws-java-sdk-1.11.319.jar each-spark-node:/usr/spark/jars
    $ scp azure-storage-7.0.0.jar each-spark-node:/usr/spark/jars 
  3. Copy the magi jar file to the spark master node.
    $ scp target/scala-2.11/magi_2.11-1.2.jar spark-master-node:~/
    

Bucket synchronization

  • Function: replicate/synchronize objects between two buckets sitting in S3, Azure or NetApp StorageGrid.
  • Options
    • --originBucket bucket, specify the origin bucket
    • --destBucket bucket, specify the destination bucket
    • --delete, delete objects in the origin bucket after the synchronization
    • --dryrun, skip the copy process
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
    • --verbose, log all missing objects in output
  1. Configure cloud endpoints and credentials in spark configuration file.

    vim /usr/spark/conf/spark-defaults.conf
    // Example 1. Use Azure as the origin
    + spark.sb.origin azure
    + spark.sb.origin.account storageaccount
    + spark.sb.origin.secretkey secretkey
    // Example 2. Use S3 as the destination
    + spark.sb.destination s3
    + spark.sb.dest.account s3account
    + spark.sb.dest.secretkey s3secretkey
    + spark.sb.dest.region us-east-1
    // Example 3. Use NetApp StorageGrid as the origin
    + spark.sb.origin storagegrid
    + spark.sb.origin.account storagegridaccount
    + spark.sb.origin.secretkey storagegridsecretkey
    + spark.sb.origin.endpointURL https://storagegridhostname:port
    

    NOTE: For AWS S3, we need to specify which region the bucket resides, by setting a value for the property "spark.sb.[origin/dest].region". For NetApp StorageGrid, we need to specify the endpoint URL, by setting a value for the property "spark.sb.[origin/dest].endpointURL".

  2. Submit the job

    $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class SyncBucket ~/magi_2.11-1.2.jar --originBucket originbucket \
    --destBucket destinationBucket > bucket-sync.log
    

Utility Functions

We also implemented three other functions that can help for testing purposes. They share the same set of configuration parameters.

Configuration

  1. Configuration for Azure
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud azure
    + spark.cloud.account storageaccount
    + spark.cloud.secretkey secretkey
    
  2. Configuration for AWS S3
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud s3
    + spark.cloud.account s3account
    + spark.cloud.secretkey s3secretkey
    + spark.cloud.region us-east-1
    
  3. Configuration for NetApp StorageGrid
    vim /usr/spark/conf/spark-defaults.conf
    + spark.cloud storagegrid
    + spark.cloud.account storagegridaccount
    + spark.cloud.secretkey storagegridsecretkey
    + spark.cloud.endpointURL https://storagegridhostname:port
    

Empty a bucket

  • Function: delete all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to delete objects
    • --dryrun, run the program without actually issuing delete operations
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class EmptyBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-empty.log
    

List a bucket

  • Function: list all objects in a bucket

  • Options

    • --bucket bucket, specify the bucket to list
    • --summary, only print the total number of ojects
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class ListBucket ~/magi_2.11-1.2.jar --bucket bucket > bucket-list.log
    

Fill a bucket

  • Function: fill a bucket with objects. The last 4 bytes of each object is the 32-bit CRC checksum calculated based on the previous bytes of the object.

  • Options

    • --bucket bucket, specify the bucket to populate
    • --size size, specify object size in bytes
    • --count count, specify the number of objects to populate
    • --prefix [numeric|letters|hex|alphanumerics|all], specify which set of prefixes to use
  • Submit the job

     $ /usr/spark/bin/spark-submit --master spark://spark-master-node-ip:7077 \
    --class FillBucket ~/magi_2.11-1.2.jar --bucket bucket \
    --size size --count 1000 > bucket-fill.log
    

Questions?

While this is not a product from NetApp, we do welcome feedbacks. Please contact ng-magi@netapp.com for feedbacks and questions.

About

High speed replication to Azure & AWS from StorageGRID

Topics

Resources

Stars

3 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages