ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ImprovingJavaIntegrationPerformance

Martin Prout edited this page Apr 28, 2017 · 5 revisions

Improving Java Integration Performance

Calling Ruby from Java involves a few challenges for JRuby:

  • Objects may need to be coerced to or from their Java equivalents, like Ruby String to Java String
  • The target method may have multiple overloads from which to pick
  • Java's reflection subsystem may introduce overhead into the calls
  • JRuby 1.7 and below will attempt to cache object wrappers, and this caching introduces overhead

This page provides some tips on how to improve the performance of calling Java methods from Ruby.

Avoiding Java Integration altogether

The "nuclear option" to improve Java Integration performance is to not use it. Calls from Ruby to Ruby have less overhead than calls from Ruby to Java, and if you wrap a Java library in a purpose-built JRuby extension (itself probably written in Java to JRuby's extension API), that will also avoid Java Integration overhead. However you can, by using the tips below, achieve the same level of performance with Java Integration in JRuby 1.7 and higher on Java 7 (where invokedynamic helps a lot).

The rest of this article will focus on improving Java Integration perf, rather than avoiding it.

Pre-coerce values used repeatedly

If calling methods from Ruby with the same coercible Ruby values, for example the same Ruby String repeatedly, pre-coercing the Ruby String to a Java String will avoid that coercion happening for every call.

str = "my string"
obj = StringReceiver.new
obj.receiveString(str) # must coerce str every time
str_java = str.to_java # pre-coerce to have a Java String in hand
obj.receiveString(str_java) # no coercion, value goes straight in

Provide signature-specific aliases for overloaded methods

When a target Java method has multiple overloads, JRuby must choose among them for every call. JRuby caches matches it has seen, but there's still a lookup process which adds overhead. Also, newer features of JRuby like its use of invokedynamic do not yet optimize overloaded Java calls.

You can use java_alias to create signature-specific aliases of a given method, avoiding the lookup overhead and allowing invokedynamic to optimize the calls.

# Our Java class, org.jruby.util.ByteList has three overloads for #append.
#
# append(int)
# append(byte[]) <== We want this one
# append(ByteList)
java_import org.jruby.util.ByteList
target = ByteList.new foo_bytes = ByteList.plain("foo") # 'foo' byte[]
target.append(foo_bytes) # must choose from the three signatures every time
class ByteList
# Add an alias for the byte[] overload
java_alias :append_bytes, :append, [Java::byte[]]
end
target.append_bytes(foo_bytes) # direct call, no lookup, optimizable

Avoid the proxy cache

As described on the Persistence page, JRuby versions 1.7 and earlier attempt to cache the wrappers we put around Java objects. This "persistence" of wrappers has a very high cost, and will be lazy or opt-in by default in JRuby versions after 1.7. By turning off wrapper persistence (via -Xji.objectProxyCache=false), Java calls that send or receive Java objects will not have the additional overhead of caching their wrappers.

Putting it all together

Here's a simple benchmark of appending a byte[] to a ByteList via a Java integration call, and the equivalent logic in the String#<< method. First, the benchmark:

java_import org.jruby.util.ByteList
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append(foo) } }
}
puts "Measure string appends (via normal Ruby)"
5.times { puts Benchmark.measure { str = "" foo = "foo" 1000000.times { str << foo } } }

The results with none of the above tips in play:

Measure bytelist appends (via Java integration)
0.285000 0.000000 0.285000 ( 0.285000)
0.177000 0.000000 0.177000 ( 0.177000)
0.187000 0.000000 0.187000 ( 0.187000)
0.168000 0.000000 0.168000 ( 0.168000)
0.172000 0.000000 0.172000 ( 0.172000)
Measure string appends (via normal Ruby)
0.113000 0.000000 0.113000 ( 0.113000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)
0.075000 0.000000 0.075000 ( 0.075000)
0.076000 0.000000 0.076000 ( 0.076000)

As you can see, the Java call has significantly more overhead, even though the resulting work done is nearly identical in both cases.

Now, we modify the benchmark according to the tips above and run with -Xji.objectProxyCache=false:

class ByteList
java_alias :append_bytes, :append, [Java::byte[]]
end
puts "Measure bytelist appends (via Java integration)"
5.times { puts Benchmark.measure { sb = ByteList.new foo = ByteList.plain("foo") 1000000.times { sb.append_bytes(foo) } }
}

The resulting numbers show that the Java call now has no overhead compared to the Ruby call. Both calls are direct and optimized well by JRuby and invokedynamic:

Measure bytelist appends (via Java integration)
0.194000 0.000000 0.194000 ( 0.193000)
0.107000 0.000000 0.107000 ( 0.107000)
0.083000 0.000000 0.083000 ( 0.083000)
0.082000 0.000000 0.082000 ( 0.082000)
0.084000 0.000000 0.084000 ( 0.084000)
Measure string appends (via normal Ruby)
0.109000 0.000000 0.109000 ( 0.108000)
0.107000 0.000000 0.107000 ( 0.107000)
0.081000 0.000000 0.081000 ( 0.081000)
0.080000 0.000000 0.080000 ( 0.080000)
0.082000 0.000000 0.082000 ( 0.082000)

(Note: Minor variation in these numbers is expected; the important detail is that both runs are now roughly equivalent speed)

Clone this wiki locally