Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo by WaldenLee2005 · Pull Request #1 · agentenv/miles · GitHub
Skip to content

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo - #1

Merged
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation
Aug 27, 2026
Merged

feat(sao): add end-to-end Terminal-Bench SAO with streaming DiLoCo#1
WaldenLee2005 merged 40 commits into
mainfrom
feat/sao-tbench21-e2e-validation

Conversation

@WaldenLee2005

@WaldenLee2005WaldenLee2005 commented Aug 27, 2026

Copy link
Copy Markdown

Summary

  • Add hash-bound offline SAO critic value pretraining and mixed-label explained-variance evaluation.
  • Implement full-parameter online SAO using DIS, observation-skipping length-adaptive GAE, and separate actor/critic optimization.
  • Connect actor and critic to independent streaming-DiLoCo synchronizers with FP32 updates, fixed quorum, one-GPU islands, and SGLang policy publication.
  • Add learned trajectory compaction and Codex/OpenEnv Terminal-Bench 2.1 rollouts with managed DinD lifecycle, native verifier evidence, HMAC validation, capacity preflights, and cleanup gates.
  • Add checkpoint-backed held-out evaluation with publication proof and pass@1-pass@4 reporting.
  • Document the validated architecture, provenance, results, operating procedure, and limitations.

Production validation

Validated Qwen3.5-0.8B on one 8xH200 node as eight one-GPU islands:

  • 356/356 signed all-task baseline rollouts accepted across all 89 Terminal-Bench 2.1 tasks.
  • 176/176 online-training trajectories accepted, including 42 learned compactions.
  • Every island completed one actor update and two critic updates without OOM.
  • Independent actor and critic streaming-DiLoCo sessions reached 8/8 quorum.
  • All islands agreed on the final actor policy hash.
  • Critic pretraining loss decreased from 5.8320 to 1.0629.
  • Corrected mixed-label critic EV was +0.1365 trajectory-weighted and +0.1353 token-weighted, with 7/8 positive batches.
  • Checkpoint-backed evaluation proved 248/309 served language tensors changed from the base policy.

Matched held-out evaluation showed no aggregate policy improvement:

MetricBaseTrained
Successful rollouts1/1801/180
pass@10.5556%0.5556%
pass@21.1111%1.1111%
pass@31.6667%1.6667%
pass@42.2222%2.2222%

The successful task changed, confirming behavioral change, but this is architecture validation rather than a learning-quality result.

Verification

  • 35/35 production-container tests passed for the corrected evaluation path.
  • 192 focused tests passed after reconciling with current main.
  • Black, isort, Ruff, Python compilation, and git diff --check passed.
  • Current main is contained in this branch; there are no merge conflicts.

Caveats

  • The upstream-integrated tree has not yet been rerun end to end on GPUs; a short production-container GPU smoke remains the next runtime gate.
  • All 176 online terminal rewards were zero and each island performed only one actor step.
  • Validation was single-node; physical multi-node Ethernet behavior was not exercised.
  • The 0.8B experiment did not demonstrate benchmark improvement.

Companion Yeto PR: agentenv/yeto#46.

Artifacts

AlexEisieand others added 30 commits July 30, 2026 11:30
Expose deterministic router, rollout-engine, and trainer process-group ports for co-resident Miles drivers. Preserve the requested attention backend when constructing LoRA models through Megatron Bridge.
Allow an external policy synchronizer to keep the rollout loop running until it returns a stop result. Always publish the final trainer weights before stopping and preserve the bounded native loop when the option is disabled.\n\nCover stop-after-publication ordering and repeated trainable-state applies that preserve optimizer state while advancing scheduler progress.
Wake an offloaded training actor before external policy synchronization initializes, keep it resident through the initial rollout weight publication, and offload it only after that publication completes. This lets initial-adapter export and application use live Megatron tensors without leaving the actor awake for the first training step.
Extend the external policy sync regression test to assert the exact onload, initialize, publish, offload, train, synchronize, and republish ordering.
Materialize bridge-converted LoRA weights on CPU while the trainer is still resident, before colocated offload releases its model storage. Consume that staged snapshot after rollout weights return, copying only the flattened adapter into a fresh CUDA allocation for CUDA IPC.
This avoids both stale offload-backed tensors and CPU file-descriptor serialization across Ray and SGLang process authentication boundaries. Add focused ordering, staging, and IPC transport regressions.
@WaldenLee2005
WaldenLee2005 merged commit 3c5e19b into mainAug 27, 2026
3 of 5 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@WaldenLee2005@AlexEisie@shouc