feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: Reduce streaming DNS failures with stale-fallback DNS cache - #323

Draft
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34
Draft

feat: Reduce streaming DNS failures with stale-fallback DNS cache#323
abelonogov-ld wants to merge 6 commits into
mainfrom
andrey/dns-sdk34

Conversation

@abelonogov-ld

@abelonogov-ldabelonogov-ld commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add CachingDns, a thread-safe OkHttp Dns wrapper that caches successful lookups (10-min TTL) and returns stale cached addresses when a fresh resolution fails — preventing UnknownHostException from killing the stream during network transitions
  • Share a ConnectionPool and the DNS resolver across StreamingDataSource restarts (context switches, foreground/background toggles, network changes) via StreamingDataSourceBuilderImpl, so cached state survives data source recreation
  • Explicitly set retryOnConnectionFailure(true) on the streaming OkHttpClient on all API levels

Background

We observed a high percentage of DNS failures in streaming connections. The root cause is that StreamingDataSource creates a new OkHttpClient on every start() call, and ConnectivityManager restarts the data source on every network change — exactly when DNS is most fragile. OkHttp uses Dns.SYSTEM (InetAddress.getAllByName) with no caching, and Android's system DNS cache has very short TTLs (sometimes ~2 seconds) that get cleared on network transitions.

This pattern of application-level DNS caching with stale fallback is well-established: Alibaba's HTTPDNS SDK, gRPC-Java's DnsNameResolver, and Square's own DnsOverHttps module all implement similar approaches. Google validated the pattern by adding DnsOptions.StaleDnsOptions to the Android framework in API 34.

Test plan

  • Unit tests for CachingDns: fresh resolution, cache hits within TTL, TTL expiry refresh, stale fallback on failure, cold failure propagation, per-hostname isolation, expiration boundary
  • Existing StreamingDataSourceTest passes (builder creates data source with new constructor args transparently)
  • Verify via logs that CachingDns warns on stale fallback and that stream reconnects succeed during network changes

Note

Medium Risk
Touches streaming network connection setup and DNS resolution behavior; incorrect caching/pooling could cause connectivity regressions or use stale IPs longer than intended.

Overview
Adds CachingDns, an OkHttp Dns wrapper that caches successful lookups with a TTL and falls back to stale cached addresses when fresh resolution fails, reducing UnknownHostException disruptions during mobile network transitions.

Updates streaming to reuse a shared DNS resolver and ConnectionPool across StreamingDataSource restarts, and wires these into the EventSource OkHttp client configuration (including enabling retryOnConnectionFailure(true)). Includes unit tests covering cache hit/expiry behavior, stale fallback, per-host caching, and eviction behavior when exceeding MAX_ENTRIES.

Written by Cursor Bugbot for commit 0a6a23e. This will update automatically on new commits. Configure here.

@abelonogov-ld
abelonogov-ld requested a review from a team as a code ownerFebruary 21, 2026 00:14

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

@abelonogov-ldabelonogov-ld changed the title feat: Reduce streaming DNS failures with stale-fallback DNS cache (API < 34) Body:feat: Reduce streaming DNS failures with stale-fallback DNS cacheFeb 21, 2026

@tanderson-ldtanderson-ld left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this needs more discussion in slack channel so we're on all the same page about possible impact and how we'll confirm there is no negative impact.

clientBuilder.readTimeout(READ_TIMEOUT_MS, TimeUnit.MILLISECONDS);
clientBuilder.dns(dns);
clientBuilder.connectionPool(connectionPool);
clientBuilder.retryOnConnectionFailure(true);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The streaming data source already follows an exponential backoff at the event source layer. We shouldn't add more try logic as the exponential backoff was specified with agreement from flag delivery to have a known load in the case of cloud outage + recovery.


if (entry != null && !entry.isExpired(now)) {
return entry.addresses;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should only use the cache in cases where the delegate fails to do lookup to minimize the impact of these changes on existing cases. My understanding of this PR is to handle a case that is not handled well today and not to improve the performance of caching in general.

Another reason is this cache has no other signals that can remove elements from it besides time, such as registered IP changes or VPN region change. It is feasible that a delegate could be aware of extra signals, but by using our own cache, we do not give the delegate an opportunity to use its possibly more complex behavior.

@abelonogov-ld
abelonogov-ld marked this pull request as draft June 3, 2026 17:40
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@abelonogov-ld@tanderson-ld