Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Add plain train.sbatch and best_train.sbatch launchers by eugenevinitsky · Pull Request #523 · Emerge-Lab/PufferDrive · GitHub
Skip to content

Add plain train.sbatch and best_train.sbatch launchers - #523

Open
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch
Open

Add plain train.sbatch and best_train.sbatch launchers#523
eugenevinitsky wants to merge 2 commits into
3.0from
ev/simple-train-sbatch

Conversation

@eugenevinitsky

Copy link
Copy Markdown

What

Adds a copy-pasteable single-GPU cluster path alongside the submitit pipeline:

  • scripts/train.sbatch — plain default training job (1 GPU, 16 CPU, 96 GB, 48 h). TRAIN_CMD is a bash array in the exact puffer train CLI dash-flag format, so it copy-pastes to and from a console command.
  • scripts/best_train.sbatch — same skeleton (192 GB, 30 h) carrying the best-known hyperparameters translated from scripts/cluster_configs/nightly_best.yaml, one flag per line with the yaml's section comments preserved.
  • Both echo a launch record (hostname, date, git commit, full command) into the job log before running, and support an optional singularity wrap when SINGULARITY_IMAGE + SINGULARITY_OVERLAY are set (required on NYU Greene). The wrap cds into the repo before sourcing the venv and %q-quotes every argument so the command survives the bash -c round-trip.
  • docs/cluster_training.md gains a short "Simple path — plain sbatch" section and a pointer from the submit_cluster.py section. The submit snippet includes the one-time mkdir -p slurm_logs — slurmd opens the stdout file before the job script runs, so the directory must exist at submit time.
  • slurm_logs/ added to .gitignore.
  • AGENTIC_PR poem per repo convention.

Why

The only documented cluster path was the ~500-line submitit pipeline, whose yaml/underscore and sweep-override formats can't be copy-pasted to or from the real puffer train CLI, and whose launched args end up buried in submitit stdout. These two files give a plain sbatch default where the exact command and code version are greppable from the job log.

Notes

Verification performed

  • bash -n on both scripts.
  • Both TRAIN_CMD flag sets parsed through pufferlib.pufferl.load_config (pufferlib resolved in-worktree) with zero unrecognized arguments; spot values match nightly_best.yaml (num_agents=4096, backbone_num_layers=3, total_timesteps=10B, reward_conditioning=False, num_goals=3).
  • Singularity branch dry-run with a stub singularity and SEED="1 2": payload arrives cwd-first, venv-sourced, with whitespace-containing args intact (--train.seed 1\ 2).
  • Non-singularity branch dry-run with a stub puffer and fake venv: full launch record echoed, command invoked verbatim, exit 0.
  • Commit contains exactly 5 files, sbatch scripts mode 100755, no build artifacts.

Risks

  • best_train.sbatch duplicates nightly_best.yaml values by design ("keep in sync" header); drift is possible until one becomes canonical.
  • Jobs run from the live checkout (no code isolation) — documented in the script headers; use submit_cluster.py when isolation, sweeps, DDP, or the Greene heartbeat is needed.
  • date -Is in the launch record requires GNU date (fine on Linux clusters; fails harmlessly on macOS).

Merge-order dependencies

  • Deliberately emits no --eval.* flags: the sibling PR fix/drive-ini-remove-hardcoded-scratch-paths removes/disables the [eval.behaviors_*] / [eval.validation_replay] drive.ini sections, and such flags would become unrecognized. This PR is safe to merge before or after it.
  • Another PR touches the quick-overview block of docs/cluster_training.md (~lines 8–13 pre-change); this PR only inserts a new section above it and one pointer sentence at the submit_cluster.py heading, so conflicts should be trivial.

🤖 Generated with Claude Code

🤖 Generated with Claude Code

Eugene Vinitskyand others added 2 commits July 10, 2026 11:15
Provide a copy-pasteable single-GPU submission path next to the
submitit pipeline: TRAIN_CMD uses the exact puffer CLI dash-flag
format, the job log records hostname/date/commit/command, and an
optional singularity wrap covers NYU Greene (cwd-safe, %q-quoted).
best_train.sbatch mirrors scripts/cluster_configs/nightly_best.yaml.
Document both in docs/cluster_training.md, including the one-time
mkdir -p slurm_logs slurmd needs before the first submit, and
gitignore slurm_logs/.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@eugenevinitsky