Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - coherentpath/data_buffer: DataBuffer provides an efficient way to maintain persistable lists of data. · GitHub
Skip to content

Repository files navigation

DataBuffer

Hex.pmDocumentation

DataBuffer is a high-performance Elixir library for buffering and batch processing data. It provides automatic flushing based on size or time thresholds, making it ideal for scenarios where you need to aggregate data before processing it in bulk.

Features

  • 🚀 High Performance - Efficient in-memory buffering with ETS-backed storage
  • Automatic Flushing - Configurable size and time-based triggers
  • 🔄 Multiple Partitions - Distribute load across multiple buffer partitions
  • 📊 Telemetry Integration - Built-in observability and monitoring
  • 🛡️ Fault Tolerant - Graceful shutdown with automatic flush on termination
  • ⏱️ Backpressure Handling - Configurable timeouts and overflow protection
  • 🎲 Jitter Support - Prevent thundering herd with configurable jitter

Installation

Add data_buffer to your list of dependencies in mix.exs:

defdepsdo[{:data_buffer,"~> 0.7.1"}]end

Quick Start

1. Define Your Buffer Module

Create a module that implements the DataBuffer behaviour:

defmoduleMyApp.EventBufferdouseDataBufferdefstart_link(opts)doDataBuffer.start_link(__MODULE__,opts)end@implDataBufferdefhandle_flush(data_stream,opts)do# Process your buffered data hereevents=Enum.to_list(data_stream)# Example: Bulk insert to databaseMyApp.Repo.insert_all("events",events)# Or send to external serviceMyApp.Analytics.track_batch(events):okendend

2. Add to Your Supervision Tree

defmoduleMyApp.ApplicationdouseApplicationdefstart(_type,_args)dochildren=[# Start your buffer with configuration{MyApp.EventBuffer,name: MyApp.EventBuffer,partitions: 4,max_size: 1000,flush_interval: 5000}]opts=[strategy: :one_for_one,name: MyApp.Supervisor]Supervisor.start_link(children,opts)endend

3. Use Your Buffer

# Insert single itemsDataBuffer.insert(MyApp.EventBuffer,%{type: "click",user_id: 123})# Insert batches for better performanceevents=[%{type: "view",user_id: 123},%{type: "click",user_id: 456},%{type: "purchase",user_id: 789}]DataBuffer.insert_batch(MyApp.EventBuffer,events)# Manually trigger flush if neededDataBuffer.flush(MyApp.EventBuffer)# Check buffer statussize=DataBuffer.size(MyApp.EventBuffer)info=DataBuffer.info(MyApp.EventBuffer)

Configuration Options

OptionTypeDefaultDescription
:nameatomrequiredProcess name for the buffer
:partitionsinteger1Number of partition processes
:max_sizeinteger5000Maximum items before automatic flush
:max_size_jitterinteger0Random jitter (0 to n) added to max_size
:flush_intervalinteger10000Time in ms between automatic flushes
:flush_jitterinteger2000Random jitter (0 to n) added to flush_interval
:flush_timeoutinteger60000Timeout in ms for flush operations
:flush_metaanynilMetadata passed to handle_flush callback

Advanced Usage

Custom Flush Metadata

Pass metadata to your flush handler for context:

{MyApp.EventBuffer,name: MyApp.EventBuffer,flush_meta: %{destination: "analytics_db",priority: :high}}# In your handle_flush callback:defhandle_flush(data_stream,opts)dometa=Keyword.get(opts,:meta)size=Keyword.get(opts,:size)casemeta.destinationdo"analytics_db"->process_analytics(data_stream)"metrics_db"->process_metrics(data_stream)endend

Multiple Buffers

You can run multiple buffers for different data types:

children=[{MyApp.EventBuffer,name: MyApp.EventBuffer,max_size: 1000},{MyApp.MetricsBuffer,name: MyApp.MetricsBuffer,max_size: 5000},{MyApp.LogBuffer,name: MyApp.LogBuffer,flush_interval: 1000}]

Synchronous Operations

For testing or specific use cases, use synchronous operations:

# Wait for flush to complete and get resultsresults=DataBuffer.sync_flush(MyApp.EventBuffer)# Dump buffer contents without flushingdata=DataBuffer.dump(MyApp.EventBuffer)

Monitoring with Telemetry

DataBuffer emits telemetry events that you can hook into:

:telemetry.attach("buffer-metrics",[:data_buffer,:flush,:stop],fn_event_name,measurements,metadata,_config->Logger.info("Flushed #{metadata.size} items from #{metadata.buffer}")end,nil)

Available events:

  • [:data_buffer, :insert, :start] / [:data_buffer, :insert, :stop]
  • [:data_buffer, :flush, :start] / [:data_buffer, :flush, :stop]

Implementation Details

Architecture

DataBuffer uses a multi-process architecture for reliability and performance:

  1. Supervisor Process - Manages the buffer's child processes
  2. Partition Processes - Handle data storage and triggering flushes
  3. Flusher Processes - Execute flush operations in separate processes

Partitioning Strategy

Data is distributed across partitions using round-robin selection. Each partition:

  • Maintains its own ETS table for data storage
  • Has independent size and time triggers
  • Flushes asynchronously without blocking other partitions

Flush Triggers

Flushes are triggered when:

  1. Size threshold is reached (max_size + random jitter)
  2. Time interval expires (flush_interval + random jitter)
  3. Manual flush is called via DataBuffer.flush/2
  4. Process termination occurs (graceful shutdown)

Backpressure Handling

When a partition is full and flushing:

  • New inserts wait for the flush to complete
  • Configurable timeout prevents indefinite blocking
  • Multiple partitions help distribute load

Fault Tolerance

  • Flushes run in separate processes to isolate failures
  • Timeouts prevent stuck flush operations
  • Graceful shutdown ensures data is flushed on termination
  • Supervisor restarts failed components

Use Cases

DataBuffer is ideal for:

  • Database Write Batching - Accumulate records for bulk inserts
  • Event Aggregation - Collect events before sending to analytics services
  • Log Processing - Buffer log entries for batch processing
  • Metrics Collection - Aggregate metrics before reporting
  • API Rate Limiting - Batch API calls to respect rate limits
  • Stream Processing - Buffer streaming data for chunk processing

Performance Considerations

  • Partition Count: More partitions = better concurrency but more memory overhead
  • Buffer Size: Larger buffers = fewer flushes but more memory usage
  • Flush Timeout: Balance between reliability and throughput
  • Jitter: Prevents synchronized flushes across multiple buffers/nodes

Testing

For testing, you can use smaller thresholds and synchronous operations:

defmoduleMyApp.EventBufferTestdouseExUnit.Casesetupdo{:ok,_pid}=MyApp.EventBuffer.start_link(name: TestBuffer,max_size: 10,flush_interval: 100){:ok,buffer: TestBuffer}endtest"buffers and flushes data",%{buffer: buffer}doDataBuffer.insert(buffer,%{id: 1})DataBuffer.insert(buffer,%{id: 2})results=DataBuffer.sync_flush(buffer)assertlength(results)==1endend

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

MIT License - see LICENSE file for details

About

DataBuffer provides an efficient way to maintain persistable lists of data.

Resources

Stars

3 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages