Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Bright Data Python SDK

The official Python SDK for Bright Data APIs. Scrape any website, get SERP results, bypass bot detection and CAPTCHAs, and access 100+ ready-made datasets.

PythonLicense

Installation

pip install brightdata-sdk

Configuration

Get your API Token from the Bright Data Control Panel:

export BRIGHTDATA_API_TOKEN="your_api_token_here"

Already logged in with the CLI? The SDK works with no configuration — it automatically falls back to the credentials stored by brightdata login.

Quick Start

This SDK is async-native. A sync client is also available (see Sync Client).

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)
asyncio.run(main())

Usage Examples

Web Scraping

asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")
print(result.data)

Web Scraping Async Mode

For non-blocking web scraping, use mode="async". This triggers a request and returns a response_id, which the SDK automatically polls until results are ready:

asyncwithBrightDataClient() asclient:
# Triggers request → gets response_id → polls until readyresult=awaitclient.scrape_url(
url="https://example.com",
mode="async",
poll_interval=5, # Check every 5 secondspoll_timeout=180# Web Unlocker async can take ~2 minutes
)
print(result.data)
# Batch scraping multiple URLs concurrentlyurls= ["https://example.com", "https://example.org", "https://example.net"]
results=awaitclient.scrape_url(url=urls, mode="async", poll_timeout=180)

How it works:

  1. Sends request to /unblocker/req → returns response_id immediately
  2. Polls /unblocker/get_result?response_id=... until ready or timeout
  3. Returns the scraped data

When to use async mode:

  • Batch scraping with many URLs
  • Background processing while continuing other work

Performance note: Web Unlocker async mode typically takes ~2 minutes to complete. For faster results on single URLs, use the default sync mode (no mode parameter).

Search Engines (SERP)

Search across Google, Bing, and Yandex. Google and Bing return parsed organic results (title/url/description/position); Yandex has no upstream parser and returns raw HTML.

asyncwithBrightDataClient() asclient:
# Google — parsed resultsresult=awaitclient.search.google(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Bing — parsed results (same shape as Google)result=awaitclient.search.bing(query="python scraping", num_results=10)
foriteminresult.data:
print(item["title"], item["url"])
# Yandex — raw HTML only (parse it yourself with BeautifulSoup etc.)result=awaitclient.search.yandex(query="python scraping", num_results=10)
print(f"{len(result.raw_html)} chars of HTML")

Batch queries: pass a list of strings instead of a single string — each query runs concurrently and you get back a List[SearchResult].

SERP Async Mode

For non-blocking SERP requests, use mode="async":

asyncwithBrightDataClient() asclient:
# Non-blocking - polls for resultsresult=awaitclient.search.google(
query="python programming",
mode="async",
poll_interval=2, # Check every 2 secondspoll_timeout=30# Give up after 30 seconds
)
foriteminresult.data:
print(item['title'], item['link'])

When to use async mode:

  • Batch operations with many queries
  • Background processing while continuing other work
  • When scraping may take longer than usual

Note: Async mode uses the same zones and returns the same data structure as sync mode - no extra configuration needed!

Web Scraper API

The SDK includes ready-to-use scrapers for popular websites: Amazon, LinkedIn, Instagram, Facebook, and more.

Pattern:client.scrape.<platform>.<method>(url)

Example: Amazon

asyncwithBrightDataClient() asclient:
# Product detailsresult=awaitclient.scrape.amazon.products(url="https://amazon.com/dp/B0CRMZHDG8")
# Reviewsresult=awaitclient.scrape.amazon.reviews(url="https://amazon.com/dp/B0CRMZHDG8")
# Sellersresult=awaitclient.scrape.amazon.sellers(url="https://amazon.com/dp/B0CRMZHDG8")

Available scrapers:

  • client.scrape.amazon — products, reviews, sellers
  • client.scrape.linkedin — profiles, companies, jobs, posts
  • client.scrape.instagram — profiles, posts, comments, reels
  • client.scrape.facebook — posts, comments, reels, pages, marketplace, events
  • client.scrape.tiktok — profiles, posts by keyword/profile/url
  • client.scrape.youtube — videos, channels, comments
  • client.scrape.reddit — posts, comments, posts by keyword
  • client.scrape.pinterest — posts, profiles, posts by keyword/profile
  • client.scrape.chatgpt — send prompts, retrieve responses
  • client.scrape.perplexity — queries with AI-ranked results
  • client.scrape.digikey — electronic component products

Browser API

Cloud-hosted Chrome instances accessible via the Chrome DevTools Protocol (CDP). The SDK builds the connection URL — you drive the browser with Playwright, Puppeteer, or Selenium.

frombrightdataimportBrightDataClientfromplaywright.async_apiimportasync_playwrightclient=BrightDataClient(
browser_username="brd-customer-<id>-zone-<zone>",
browser_password="<password>",
)
url=client.browser.get_connect_url(country="us") # country is optionalasyncwithasync_playwright() aspw:
browser=awaitpw.chromium.connect_over_cdp(url)
page=awaitbrowser.new_page()
awaitpage.goto("https://example.com")
html=awaitpage.content()
awaitbrowser.close()

When to use: sites that require full browser automation — JS rendering, login flows, interactive clicks. For plain HTML fetches, prefer client.scrape_url().

Discover API

AI-ranked web search. Unlike SERP (which returns engine-ordered results), Discover takes a query plus an intent phrase and re-ranks by relevance. Optionally extracts full page content as markdown.

asyncwithBrightDataClient() asclient:
result=awaitclient.discover(
query="artificial intelligence trends 2026",
intent="latest AI technology developments",
country="us",
num_results=10,
)
foriteminresult.data:
print(f"[{item['relevance_score']:.2f}] {item['title']}{item['link']}")

For long-running discoveries, trigger and poll separately:

job=awaitclient.discover_trigger(query="...", intent="...")
result=awaitjob.wait_and_fetch(timeout=60)

When to use Discover vs SERP: Discover when you want entity-level relevance ranking driven by a natural-language intent (e.g. "find sustainability-focused AI companies"). SERP when you want raw search engine results.

Scraper Studio

Run custom collectors built in Bright Data's Scraper Studio. One call triggers the job, polls until ready, and returns the records:

asyncwithBrightDataClient() asclient:
data=awaitclient.scraper_studio.run(
collector="c_abc123", # your collector IDinput={"url": "https://example.com"}, # collector-specific inputtimeout=180,
)
print(f"Got {len(data)} records")

For manual control over the lifecycle, use client.scraper_studio.trigger().status().fetch().

Datasets API

Access 100+ ready-made datasets from Bright Data — pre-collected, structured data from popular platforms.

asyncwithBrightDataClient() asclient:
# Filter a dataset — returns a snapshot_idsnapshot_id=awaitclient.datasets.imdb_movies(
filter={"name": "title", "operator": "includes", "value": "black"},
records_limit=5
)
# Download when ready (polls until snapshot is complete)data=awaitclient.datasets.imdb_movies.download(snapshot_id)
print(f"Got {len(data)} records")
# Quick sample: .sample() auto-discovers fields, no filter needed# Works on any datasetsnapshot_id=awaitclient.datasets.imdb_movies.sample(records_limit=5)

Export results to file:

frombrightdata.datasetsimportexportexport(data, "results.json") # JSONexport(data, "results.csv") # CSVexport(data, "results.jsonl") # JSONL

Available dataset categories:

  • E-commerce: Amazon, Walmart, Shopee, Lazada, Zalando, Zara, H&M, Shein, IKEA, Sephora, and more
  • Business intelligence: ZoomInfo, PitchBook, Owler, Slintel, VentureRadar, Manta
  • Jobs & HR: Glassdoor (companies, reviews, jobs), Indeed (companies, jobs), Xing
  • Reviews: Google Maps, Yelp, G2, Trustpilot, TrustRadius
  • Social media: Pinterest (posts, profiles), Facebook Pages
  • Real estate: Zillow, Airbnb, and 8+ regional platforms
  • Luxury brands: Chanel, Dior, Prada, Balenciaga, Hermes, YSL, and more
  • Entertainment: IMDB, NBA, Goodreads

Discover available fields:

metadata=awaitclient.datasets.imdb_movies.get_metadata()
forname, fieldinmetadata.fields.items():
print(f"{name}: {field.type}")

Async Usage

Run multiple requests concurrently:

importasynciofrombrightdataimportBrightDataClientasyncdefmain():
asyncwithBrightDataClient() asclient:
urls= ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]
tasks= [client.scrape_url(url) forurlinurls]
results=awaitasyncio.gather(*tasks)
asyncio.run(main())

Manual Trigger/Poll/Fetch

For long-running scrapes:

asyncwithBrightDataClient() asclient:
# Triggerjob=awaitclient.scrape.amazon.products_trigger(url="https://amazon.com/dp/B123")
# Wait for completionawaitjob.wait(timeout=180)
# Fetch resultsdata=awaitjob.fetch()

Sync Client

For simpler use cases, use SyncBrightDataClient:

frombrightdataimportSyncBrightDataClientwithSyncBrightDataClient() asclient:
result=client.scrape_url("https://example.com")
print(result.data)
# All methods work the sameresult=client.scrape.amazon.products(url="https://amazon.com/dp/B123")
result=client.search.google(query="python")

See docs/sync_client.md for details.

Account & Zones

Inspect the account, list zones, and clean up unused ones:

asyncwithBrightDataClient() asclient:
# Verify token + connectivity (never raises — returns False on any error)ifawaitclient.test_connection():
print("Connected")
# Account status, usage stats, credit balanceinfo=awaitclient.get_account_info()
print(info)
# All zones on this account, with type and usagezones=awaitclient.list_zones()
forzinzones:
print(z["name"], z["type"])
# Delete a zone — destructive, confirm before callingawaitclient.delete_zone("zone_to_remove")

Troubleshooting

RuntimeError: SyncBrightDataClient cannot be used inside async context

# Wrong - using sync client in async functionasyncdefmain():
withSyncBrightDataClient() asclient: # Error!
...
# Correct - use async clientasyncdefmain():
asyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("https://example.com")

RuntimeError: BrightDataClient not initialized

# Wrong - forgot context managerclient=BrightDataClient()
result=awaitclient.scrape_url("...") # Error!# Correct - use context managerasyncwithBrightDataClient() asclient:
result=awaitclient.scrape_url("...")

License

MIT License

About

Bright Data's python SDK, use it to call bright data's scrape and search tools. bypass any Bot-detection or Captcha and extract data from any website in seconds.

Resources

Stars

92 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages