Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - kejunxiao/ShoppingBench · GitHub
Skip to content

Repository files navigation

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

Paper

Overview

ShoppingBench is a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products.

Features

  • various real-world shopping intents
  • a large-scale shopping sandbox
  • Comprehensive evaluation metrics

Dataset

The ShoppingBench dataset includes:

  1. documents.jsonl.gz: A compressed file containing product documents (located in resources/ directory)

    • To decompress: gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
    • Size: ~1.4GB compressed, ~4.8GB uncompressed
  2. Test files: Located in the data/ directory

    • synthesize_product_test.jsonl: Product Intent test cases
    • synthesize_shop_test.jsonl: Shop Intent test cases
    • synthesize_voucher_test.jsonl: Voucher Intent test cases
    • synthesize_web_simpleqa_test.jsonl: Web search Intent test cases

Environment Setup

prerequist

  1. install java (jdk21 recommended)

  2. install uv

  3. decompress documents.jsonl.gz to get documents.jsonl in resources folder:

    gunzip -c resources/documents.jsonl.gz > resources/documents.jsonl
  4. prepare related KEY

export OPENAI_API_KEY="your openai api key"export OPENAI_BASE_URL="your openai base url"export SERPER_KEY="your serper web search key"

Python Environment Installation and Search Engine Preparation

Run the initialization script to set up the Python environment and start the product search engine:

./init_env.sh

After running the environment setup script, the search engine will be automatically started in the background.

Running Inference and Evaluation

To run model inference on test data and evaluate the models for different intents:

  1. Run the inference scripts (take gpt-4.1 as example):

    The script will automatically create necessary directories and validate the data folder structure before running model inference and evaluation.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

    the inference process will be running in background, you can check the log in logs folder. you can uncomment the specific line to evaluate the inference result or kill the inference process.

  2. Run the evaluation scripts (take gpt-4.1 as example):

    Please update run.sh by uncommenting the line for run_evaluate.py and commenting out the line for run_rollout, then rerun the scripts.

    ./run.sh product rollout gpt-4.1
    ./run.sh shop rollout gpt-4.1
    ./run.sh voucher rollout gpt-4.1
    ./run.sh web simpleqa_rollout gpt-4.1

Training SFT and RL Models

SFT Environment Installation

Install SFT environment and its dependencies (llama factory):

cd src/sft/LLaMA-Factory
uv pip install -e ".[torch,metrics,deepspeed]" --no-build-isolation

RL Environment Installation

Install RL environment and its dependencies:

cd src/rl
USE_MEGATRON=0 bash install_vllm_sglang_mcore.sh
# verl
uv pip install -e .

Usage

To train the SFT and RL models:

For SFT training:

  1. prepare sft data

    cd src/sft
    mkdir data

    place dataset_info.json and training data in the data directory

  2. run sft scripts

 ./submit.sh yaml_file

For RL training:

  1. run rl scripts
    cd src/rl
    ./run_grpo.sh

Paper

For more details about ShoppingBench, please refer to our paper

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages