Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Bright Data Scraper Studio (Python)

A minimal Python starter for running a Bright Data Scraper Studio collector via the Data Collection API: trigger a job with a list of URLs and download the results.

Bright Data Promo

Open in CodeSandbox, sign in with GitHub, then fork the repository to begin making changes.


Table of contents


Overview

Bright Data Scraper Studio is a low-code IDE for building custom web scraping collectors on the Bright Data platform. Once a collector is published it exposes two HTTP endpoints:

StepEndpointPurpose
1POST /dca/trigger?collector=<id>Queue one or more inputs for the collector
2GET /dca/dataset?id=<snapshot_id>Download the collected data when ready

This repository wraps those two calls in about 150 lines of Python so you can copy, paste and ship.


Features

  • Trigger a Scraper Studio collector via the /dca/trigger endpoint
  • Poll /dca/dataset until results are ready
  • Env-var config via .env (no secrets in code)
  • Retry with exponential backoff for transient errors (5xx and network); fails fast on 4xx
  • Library helpers: trigger_with_url, trigger_with_urls, run_scraper
  • Saves the raw JSON response to a timestamped file

Prerequisites

  • Python 3.8 or higher
  • A Bright Data account with an API token
  • A published collector in Scraper Studio; copy its Collector ID (starts with c_)

Installation

git clone https://github.com/brightdata/bright-data-scraper-studio-python-project.git
cd bright-data-scraper-studio-python-project
pip install -r requirements.txt
cp .env.example .env # then edit .env with your token and collector ID

Dependencies

  • requests: HTTP client for the Bright Data API
  • colorama: colored terminal output
  • python-dotenv: load .env files into os.environ

Usage

python index.py

Results are written to a scraper_studio_results_<timestamp>.json file in the project directory.


Configuration

Two environment variables are required. Set them in .env, in your shell, or hardcode them in index.py:

VariableWhere to find it
BRIGHT_DATA_API_TOKENBright Data dashboard, Account Settings → API Tokens
BRIGHT_DATA_COLLECTOR_IDScraper Studio: open your collector, copy the ID from the URL (starts with c_)

You can also tune the polling and retry behavior at the top of index.py:

POLL_INTERVAL_S=5# delay between dataset checks (seconds)MAX_POLL_ATTEMPTS=60# give up after ~5 minutesMAX_RETRIES=3# for transient HTTP failures

The shape of SAMPLE_URLS must match the input schema you defined in Scraper Studio. The default sample assumes a single url field. If your collector uses different inputs (for example, keyword, zip_code, category), update the dictionaries accordingly.


How it works

 +-----------------+ POST /dca/trigger +-------------------+
| Your script | --------------------------> | Scraper Studio |
| (index.py) | <-- { collection_id } ----- | Collector |
+-----------------+ +-------------------+
| |
| GET /dca/dataset?id=<snapshot_id> |
| (poll every 5s, retry 5xx with backoff) |
| <--- [ { ...record... }, ... ] -------------- |
v
scraper_studio_results_<timestamp>.json

The script polls /dca/dataset every five seconds for up to five minutes. A non-empty JSON array is treated as a finished snapshot. Transient errors (5xx and network) are retried with exponential backoff (1s, 2s, 4s); 4xx errors fail immediately so you fix the request rather than retry it.


Examples

Run with your own URLs

Replace SAMPLE_URLS in index.py:

SAMPLE_URLS= [
{"url": "https://example.com/product/1"},
{"url": "https://example.com/product/2"},
]

Custom input schema

If your collector expects something other than url, pass whatever fields it defines:

inputs= [
{"keyword": "wireless headphones", "country": "US"},
{"keyword": "standing desk", "country": "DE"},
]
run_scraper(inputs)

Use as a library

run_scraper, trigger_with_url, trigger_with_urls and save_results are top-level functions:

fromindeximporttrigger_with_urls, save_resultsdata=trigger_with_urls([
"https://example.com/page-1",
"https://example.com/page-2",
])
save_results(data, "my_run.json")

Output

  • Results are saved as JSON files named scraper_studio_results_<ISO timestamp>.json.
  • The file contains the raw collector output: one record per input URL by default.

Sample console output

Bright Data Scraper Studio
==============================
Starting Scraper Studio collector...
Queueing 3 input(s)
Job queued. Snapshot ID: j_abc123
Polling for results...
Attempt 1/60 - building
Attempt 2/60 - building
Attempt 3/60 - building
Results downloaded.
Saved to scraper_studio_results_2026-05-22T10-30-45-123456.json
Done.

Security

Never commit your .env file. The shipped .gitignore blocks .env and .env.local.

If you accidentally commit a real BRIGHT_DATA_API_TOKEN:

  1. Rotate the token immediately at brightdata.com/cp/setting.
  2. Use git filter-repo or BFG Repo-Cleaner to remove the secret from history.
  3. Force-push and notify anyone who may have cloned the leak.

Support


License

This project is licensed under the MIT License. See LICENSE for details.

About

Bright Data Scraper Studio Python boilerplate code

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages