Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

🚚 Bulker

Bulker is a tool for streaming and batching large amount of semi-structured data into data warehouses. It uses Kafka internally

How it works?

Send and JSON object to Bulker HTTP endpoint, and it will make sure it will be saved to data warehouse:

  • JSON flattening. Your object will be flattened - {a: {b: 1}} becomes {a_b: 1}
  • Schema managenent for semi-structured data. For each fuekd, bulker will make sure that a corresponding column exist in destination table. If not, Bulker will create. Type will be best-guessed by value, or it could be explicitely set via type hint as in {"a": "test", "__sql_type_a": "varchar(4)"}
  • Reliability. Bulker will put the object to Kafka Queue immidiatly, so if datawarehouse is down, data won't be lost
  • Streaming or Batching. Bulker will send data to datawarehouse either as soon it become available in Kafka (streaming) or after some time (batching). Most data warehouses won't tolerate large number of inserts, that's why we implemented batching

Bulker is a 💜 of Jitsu, an open-source data integration platform.

See full list of features below

Bulker is also available as a go library if you want to embed it into your application as opposed to use a HTTP-server

Features

  • 🛢️ Batching - Bulker sends data in batches in most efficient way for particular database. For example, for Postgres it uses COPY command, for BigQuery it uses batch-files
  • 🚿 Streaming - alternatively, Bulker can stream data to database. It is useful when number of records is low. Up to 10 records per second for most databases
  • 🐫 Deduplication - if configured, Bulker will deduplicate records by primary key
  • 📋 Schema management - Bulker creates tables and columns on the fly. It also flattens nested JSON-objects. Example if you send {"a": {"b": 1}} to bulker, it will make sure that there is a column a_b in the table (and will create it)
  • 🦾 Implicit typing - Bulker infers types of columns from JSON-data.
  • 📌 Explicit typing - Explicit types can be by type hints that are placed in JSON. Example: for event {"a": "test", "__sql_type_a": "varchar(4)"} Bulker will make sure that there is a column a, and it's type is varchar(4).
  • 📈 Horizontal Scaling. Bulker scales horizontally. Too much data? No problem, just add Bulker instances!
  • 📦 Dockerized - Bulker is dockerized and can be deployed to any cloud provider and k8s.
  • ☁️ Cloud Native - each Bulker instance is stateless and is configured by only few environment variables.

Supported databases

Bulker supports the following databases:

  • ✅ PostgresSQL
  • ✅ Redshit
  • ✅ Snowflake
  • ✅ Clickhouse
  • ✅ BigQuery
  • ✅ MySQL
  • ✅ S3
  • ✅ GCS

Please see Compatibility Matrix to learn what Bulker features are supported by each database.

Documentation Links

Note We highly recommend to read Core Concepts below before diving into details

Core Concepts

Destinations

Bulker operates with destinations. Destination is database or storage service (e.g. S3, GCS). Each destination has an ID and configuration which is represented by JSON object.

Bulker exposes HTTP API to load data into destinations, where those destinations are referenced by their IDs.

If destination is a database, you'll need to provide a destination table name.

Event

The main unit of data in Bulker is an event. Event is a represented JSON-object

Batching and Streaming (aka Destination Mode)

Bulker can send data to database in two ways:

  • Streaming. Bulker sends evens to destinaion one by one. It is useful when number of events is low (less than 10 events per second for most DBs).
  • Batching. Bulker accumulates events in batches and sends them periodically once batch is full or timeout is reached. Batching is more efficient for large amounts of events. Especially for cloud data-warehouses (e.g. Postgres, Clickhouse, BigQuery).

Primary Keys and Deduplication

Optionally, Bulker can deduplicate events by primary key. It is useful when you same event can be sent to Bulker multiple times. If available, Bulker uses primary keys, but for some data warehouses alternative strategies are used.

Read more about deduplication »

About

Service for bulk-loading data to databases with automatic schema management (Redshift, Snowflake, BigQuery, ClickHouse, Postgres, MySQL)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages