feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelinescollects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field | type | required | description
about | text | | A concise profile summary...
connections | number | | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@karaposu