The e6data Connector for Python provides an interface for writing Python applications that can connect to e6data and perform operations. It includes automatic support for blue-green deployments, ensuring seamless failover during server updates without query interruption.
Make sure to install below dependencies and wheel before install e6data-python-connector.
# Amazon Linux / CentOS dependencies
yum install python3-devel gcc-c++ -y
# Ubuntu/Debian dependencies
apt install python3-dev g++ -y
# Windows dependencies# Install Visual C++ Build Tools from:# https://visualstudio.microsoft.com/visual-cpp-build-tools/# Select the "Desktop Development with C++" option during installation.# Pip dependencies
pip install wheelTo install the Python package, use the command below:
pip install --no-cache-dir e6data-python-connector- Open Inbound Port 80 in the Engine Cluster.
- Limit access to Port 80 according to your organizational security policy. Public access is not encouraged.
- Access Token generated in the e6data console.
Use your e6data Email ID as the username and your access token as the password.
frome6data_python_connectorimportConnection# For connection pooling (recommended for concurrent operations)frome6data_python_connectorimportConnectionPoolusername='<username>'# Your e6data Email ID.password='<password>'# Access Token generated in the e6data console.host='<host>'# IP address or hostname of the cluster to be used.database='<database>'# Database to perform the query on.port=80# Port of the e6data engine.catalog_name='<catalog_name>'# Single connection (for simple, single-threaded use)conn=Connection(
host=host,
port=port,
username=username,
database=database,
password=password
)
# Or use connection pool (for concurrent/multi-threaded use)pool=ConnectionPool(
min_size=2,
max_size=10,
host=host,
port=port,
username=username,
database=database,
password=password
)The Connection class supports the following parameters:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
host | str | Yes | - | IP address or hostname of the e6data cluster |
port | int | Yes | - | Port of the e6data engine (typically 80) |
username | str | Yes | - | Your e6data Email ID |
password | str | Yes | - | Access Token generated in the e6data console |
database | str | No | None | Database to perform queries on |
catalog | str | No | None | Catalog name |
cluster_name | str | No | None | Name of the cluster for cluster-specific operations |
secure | bool | No | False | Enable SSL/TLS for secure connections |
ssl_cert | str/bytes | No | None | Path to CA certificate (PEM) or certificate bytes for HTTPS connections |
auto_resume | bool | No | True | Automatically resume cluster if suspended |
grpc_options | dict | No | None | Additional gRPC configuration options |
debug | bool | No | False | Enable debug logging for troubleshooting |
require_fastbinary | bool | No | True | Require fastbinary module for Thrift deserialization. Set to False to use pure Python implementation if system dependencies cannot be installed |
To establish a secure connection using SSL/TLS:
conn=Connection(
host=host,
port=443, # Typically 443 for secure connectionsusername=username,
password=password,
database=database,
cluster_name='production-cluster',
secure=True# Enable SSL/TLS
)When working with multiple clusters, specify the cluster name:
conn=Connection(
host=host,
port=port,
username=username,
password=password,
database=database,
cluster_name='analytics-cluster-01', # Specify cluster namesecure=True
)When connecting through HAProxy with HTTPS, you can provide a custom CA certificate for secure connections. The ssl_cert parameter accepts either a file path to a PEM certificate or the certificate content as bytes.
Using a CA certificate file path:
conn=Connection(
host=host,
port=443,
username=username,
password=password,
database=database,
secure=True,
ssl_cert='/path/to/ca-cert.pem'# Path to your CA certificate
)Reading certificate content as bytes:
# Read certificate file and pass as byteswithopen('/path/to/ca-cert.pem', 'rb') ascert_file:
cert_data=cert_file.read()
conn=Connection(
host=host,
port=443,
username=username,
password=password,
database=database,
secure=True,
ssl_cert=cert_data# Certificate content as bytes
)Using system CA bundle for publicly signed certificates:
# When ssl_cert is None, system default CA bundle is usedconn=Connection(
host=host,
port=443,
username=username,
password=password,
database=database,
secure=True# Uses system CA bundle by default
)Connection pooling with custom CA certificate:
pool=ConnectionPool(
min_size=2,
max_size=10,
host=host,
port=443,
username=username,
password=password,
database=database,
secure=True,
ssl_cert='/path/to/ca-cert.pem'# Custom CA certificate for pool connections
)The e6data connector uses the fastbinary module (from Apache Thrift) for optimal performance when deserializing data. This module requires system-level dependencies (python3-devel and gcc-c++) to be installed.
Default Behavior (Recommended):
By default, the connector requires fastbinary to be available. If it's not found, the connection will fail immediately with a clear error message:
conn=Connection(
host=host,
port=port,
username=username,
password=password,
database=database
)
# Raises exception if fastbinary is not availableFallback to Pure Python:
If you cannot install system dependencies (e.g., in restricted environments, serverless platforms, or containers without build tools), you can disable the fastbinary requirement. The connector will fall back to a pure Python implementation with a performance penalty:
conn=Connection(
host=host,
port=port,
username=username,
password=password,
database=database,
require_fastbinary=False# Allow operation without fastbinary
)
# Logs warning but continues with pure Python implementationWhen to use require_fastbinary=False:
- Running in AWS Lambda or other serverless environments
- Docker containers built without compilation tools
- Restricted environments where system packages cannot be installed
- Development/testing environments where performance is not critical
Performance Impact:
- With
fastbinary: Optimal performance for data deserialization - Without
fastbinary(pure Python): ~2-3x slower deserialization, but otherwise fully functional
Note: It's strongly recommended to install system dependencies when possible for best performance. The require_fastbinary=False option should only be used when system dependencies cannot be installed.
query='SELECT * FROM <TABLE_NAME>'# Replace with the query.cursor=conn.cursor(catalog_name=catalog_name)
query_id=cursor.execute(query) # The execute function returns a unique query ID, which can be use to abort the query.all_records=cursor.fetchall()
forrowinall_records:
print(row)To fetch all the records:
records=cursor.fetchall()To fetch one record:
record=cursor.fetchone()To fetch limited records:
limit=500records=cursor.fetchmany(limit)To fetch all the records in buffer to reduce memory consumption:
records_iterator=cursor.fetchall_buffer() # Returns generatorforiteminrecords_iterator:
print(item)To get the execution plan after query execution:
importjsonexplain_response=cursor.explain_analyse()
query_planner=json.loads(explain_response.get('planner'))To abort a running query:
query_id='<query_id>'# query id from execute function response.cursor.cancel(query_id)To get the status of a running or completed query:
query_id='<query_id>'# query id from execute function response.status_response=cursor.status(query_id)
# Check if query execution is completeis_complete=status_response.status# Returns True when query is done, False if still running# Get the total row countrow_count=status_response.rowCount# Total number of rows in the result setprint(f"Query complete: {is_complete}")
print(f"Row count: {row_count}")The status() method is useful for:
- Monitoring long-running queries: Poll the status periodically to check if execution is complete
- Checking row counts: Get the total number of rows without fetching all results
- Query progress tracking: Integrate with monitoring systems or progress bars
- Conditional fetching: Decide whether to fetch results based on completion status
Example - Polling for query completion:
importtimecursor=conn.cursor(catalog_name=catalog_name)
query_id=cursor.execute("SELECT * FROM large_table")
# Poll until query is completewhileTrue:
status_response=cursor.status(query_id)
ifstatus_response.status:
print(f"Query complete! Total rows: {status_response.rowCount}")
breakelse:
print("Query still running...")
time.sleep(1) # Wait 1 second before checking again# Now fetch the resultsresults=cursor.fetchall()Switch database in an existing connection:
database='<new_database_name>'# Replace with the new database.cursor=conn.cursor(database, catalog_name)importjsonquery='SELECT * FROM <TABLE_NAME>'cursor=conn.cursor(catalog_name)
query_id=cursor.execute(query) # execute function returns query id, can be use for aborting the query.all_records=cursor.fetchall()
explain_response=cursor.explain_analyse()
query_planner=json.loads(explain_response.get('planner'))
execution_time=query_planner.get("total_query_time") # In millisecondsqueue_time=query_planner.get("executionQueueingTime") # In millisecondsparsing_time=query_planner.get("parsingTime") # In millisecondsrow_count=query_planner.rowcountThe following code returns a dictionary of all databases, all tables and all columns connected to the cluster currently in use. This function can be used without passing database name to get list of all databases.
databases=conn.get_schema_names() # To get list of databases.print(databases)
database='<database_name>'# Replace with actual database name.tables=conn.get_tables(database=database) # To get list of tables from a database.print(tables)
table_name='<table_name>'# Replace with actual table name.columns=conn.get_tables(database=database, table=table_name) # To get the list of columns from a table.columns_with_type=list()
"""Getting the column name and type."""forcolumnincolumns:
columns_with_type.append(dict(column_name=column.fieldName, column_type=column.fieldType))
print(columns_with_type)It is recommended to clear the cursor, close the cursor and close the connection after running a function as a best practice. This enhances performance by clearing old data from memory.
cursor.clear() # Not needed when aborting a querycursor.close()
conn.close()The following code is an example which combines a few functions described above.
frome6data_python_connectorimportConnectionimportjsonusername='<username>'# Your e6data Email ID.password='<password>'# Access Token generated in the e6data console.host='<host>'# IP address or hostname of the cluster to be used.database='<database>'# # Database to perform the query on.port=80# Port of the e6data engine.sql_query='SELECT * FROM <TABLE_NAME>'# Replace with the actual query.catalog_name='<catalog_name>'# Replace with the actual catalog name.conn=Connection(
host=host,
port=port,
username=username,
database=database,
password=password
)
cursor=conn.cursor(db_name=database, catalog_name=catalog_name)
query_id=cursor.execute(sql_query)
all_records=cursor.fetchall()
explain_response=cursor.explain_analyse()
planner_result=json.loads(explain_response.get('planner'))
execution_time=planner_result.get("total_query_time") /1000# Converting into seconds.row_count=cursor.rowcountcolumns= [col[0] forcolincursor.description] # Get the column names and merge them with the results.results= []
forrowinall_records:
row=dict(zip(columns, row))
results.append(row)
print(row)
print('Total row count {}, Execution Time (seconds): {}'.format(row_count, execution_time))
cursor.clear()
cursor.close()
conn.close()The e6data Python Connector provides automatic zero downtime deployment support through intelligent blue-green deployment strategy management:
Your existing applications automatically benefit from zero downtime deployment without any modifications:
# Your existing code works exactly the samefrome6data_python_connectorimportConnectionconn=Connection(
host='your-host',
port=80,
username='your-email',
password='your-token',
database='your-database'
)
cursor=conn.cursor()
cursor.execute("SELECT * FROM your_table")
results=cursor.fetchall()- Detects active deployment strategy (blue/green) on connection
- Caches strategy information for optimal performance
- Automatically switches strategies when deployments occur
- Running queries continue uninterrupted during deployments
- New queries automatically use the new deployment strategy
- Graceful transitions ensure no query loss or failures
- < 100ms additional latency on first connection (one-time cost)
- 0ms overhead for 95% of queries (cached strategy)
- < 1KB additional memory usage per connection
- Full support for multi-threaded applications
- Process-safe shared memory management
- Concurrent query execution without conflicts
For enhanced monitoring and performance tuning:
# Enhanced gRPC configuration for zero downtimegrpc_options= {
'keepalive_timeout_ms': 60000, # 1 minute keepalive timeout'keepalive_time_ms': 30000, # 30 seconds keepalive interval'max_receive_message_length': 100*1024*1024, # 100MB'max_send_message_length': 100*1024*1024, # 100MB
}
conn=Connection(
host='your-host',
port=80,
username='your-email',
password='your-token',
database='your-database',
grpc_options=grpc_options
)Configure zero downtime features using environment variables:
# Strategy cache timeout (default: 300 seconds)export E6DATA_STRATEGY_CACHE_TIMEOUT=300
# Maximum retry attempts (default: 5)export E6DATA_MAX_RETRY_ATTEMPTS=5
# Enable debug logging for strategy operationsexport E6DATA_STRATEGY_LOG_LEVEL=INFOUse the included mock server for testing and development:
# Terminal 1: Start mock server
python mock_grpc_server.py
# Terminal 2: Run test client
python test_mock_server.py
# Or use the convenience script
./run_mock_test.shExplore detailed documentation in the docs/zero-downtime/ directory:
- 📋 Overview - Complete guide and feature overview
- 🔧 API Reference - Detailed API documentation
- 🌊 Flow Documentation - Process flows and diagrams
- 💼 Business Logic - Business rules and decisions
- 🏗️ Architecture - System architecture and design
- ⚙️ Configuration - Complete configuration guide
- 🧪 Testing - Testing strategies and tools
- 🔍 Troubleshooting - Common issues and solutions
- 🚀 Migration Guide - Step-by-step migration instructions
| Feature | Benefit |
|---|---|
| Zero Downtime | Applications continue running during e6data deployments |
| Automatic | No code changes or manual intervention required |
| Reliable | Robust error handling and automatic recovery |
| Fast | Minimal performance impact with intelligent caching |
| Safe | Thread-safe and process-safe operation |
| Monitored | Comprehensive logging and monitoring capabilities |
Existing applications automatically benefit from zero downtime deployment:
- Update connector:
pip install --upgrade e6data-python-connector - No code changes: Your existing code works without modifications
- Monitor: Use enhanced logging to monitor strategy transitions
- Validate: Test with your existing applications
For detailed migration instructions, see the Migration Guide.
- Use
fetchall_buffer()for memory-efficient large result sets - Automatic cleanup of query-strategy mappings
- Bounded memory usage with TTL-based caching
- Configure gRPC options for optimal network performance
- Intelligent keepalive settings for connection stability
- Message size optimization for large queries
The e6data Python connector now includes a built-in connection pool for efficient connection management and reuse across multiple threads. The ConnectionPool class provides:
- Thread-safe connection reuse: Each thread automatically reuses its assigned connection
- Automatic lifecycle management: Handles connection creation, health checks, and cleanup
- Overflow connections: Creates temporary connections when pool is exhausted
- Connection health monitoring: Automatic detection and replacement of broken connections
- Statistics tracking: Monitor pool usage and performance
frome6data_python_connectorimportConnectionPool# Create a connection poolpool=ConnectionPool(
min_size=2, # Minimum connections to maintainmax_size=10, # Maximum connections in poolmax_overflow=5, # Additional temporary connections allowedtimeout=30.0, # Timeout for getting connection (seconds)recycle=3600, # Maximum age before recycling (seconds)debug=False, # Enable debug loggingpre_ping=True, # Check connection health before use# Connection parametershost=host,
port=port,
username=username,
password=password,
database=database,
catalog=catalog_name,
cluster_name=cluster_name,
secure=True
)
# Get connection and execute queryconn=pool.get_connection()
cursor=conn.cursor()
cursor.execute("SELECT * FROM table")
results=cursor.fetchall()
# Return connection to pool (important!)pool.return_connection(conn)
# Clean up when donepool.close_all()The context manager pattern ensures connections are automatically returned to the pool:
frome6data_python_connectorimportConnectionPoolpool=ConnectionPool(
min_size=2,
max_size=10,
host=host,
port=port,
username=username,
password=password,
database=database
)
# Connection automatically returned to pool after usewithpool.get_connection_context() asconn:
cursor=conn.cursor()
cursor.execute("SELECT * FROM table")
results=cursor.fetchall()
print(results)Connection pooling is especially beneficial for concurrent query execution:
importconcurrent.futuresfrome6data_python_connectorimportConnectionPooldefexecute_query(pool, query_id, query):
"""Execute a query using a pooled connection."""# Each thread will reuse its assigned connectionconn=pool.get_connection()
try:
cursor=conn.cursor()
cursor.execute(query)
results=cursor.fetchall()
returnf"Query {query_id}: {len(results)} rows"finally:
pool.return_connection(conn)
# Create poolpool=ConnectionPool(
min_size=3,
max_size=10,
host=host,
port=port,
username=username,
password=password,
database=database
)
# Execute multiple queries concurrentlyqueries= [
"SELECT COUNT(*) FROM table1",
"SELECT AVG(value) FROM table2",
"SELECT MAX(date) FROM table3"
]
withconcurrent.futures.ThreadPoolExecutor(max_workers=5) asexecutor:
futures= [
executor.submit(execute_query, pool, i, query)
fori, queryinenumerate(queries)
]
forfutureinconcurrent.futures.as_completed(futures):
print(future.result())
# Clean uppool.close_all()| Parameter | Type | Default | Description |
|---|---|---|---|
min_size | int | 2 | Minimum number of connections to maintain |
max_size | int | 10 | Maximum number of connections in pool |
max_overflow | int | 5 | Additional temporary connections allowed |
timeout | float | 30.0 | Timeout for getting connection (seconds) |
recycle | int | 3600 | Maximum connection age before recycling (seconds) |
debug | bool | False | Enable debug logging for pool operations |
pre_ping | bool | True | Check connection health before returning from pool |
# Get pool statisticsstats=pool.get_statistics()
print(f"Active connections: {stats['active_connections']}")
print(f"Idle connections: {stats['idle_connections']}")
print(f"Total requests: {stats['total_requests']}")
print(f"Failed connections: {stats['failed_connections']}")Connection pooling is recommended when:
- Executing multiple queries concurrently
- Building web applications or APIs
- Running batch processing jobs
- Reducing connection overhead
- Improving application performance
For simple, single-threaded applications, you can still use direct connections:
frome6data_python_connectorimportConnectionconn=Connection(
host=host,
port=port,
username=username,
password=password,
database=database
)
cursor=conn.cursor()
cursor.execute("SELECT * FROM table")
results=cursor.fetchall()
conn.close()- Automatic connection health monitoring
- Graceful connection recovery and retry logic
- Blue-green deployment support with automatic failover
Enable comprehensive debugging to troubleshoot connection and query issues:
frome6data_python_connectorimportConnectionconn=Connection(
host=host,
port=port,
username=username,
password=password,
database=database,
debug=True# Enable debug logging
)When debug=True, the following features are enabled:
- Python logging at DEBUG level for all operations
- Blue-green strategy transition logging
- Connection lifecycle logging
- Query execution detailed logging
For low-level gRPC network debugging (HTTP/2 frames, TCP events), set environment variables before running your Python script:
# Enable gRPC network tracingexport GRPC_VERBOSITY=DEBUG
export GRPC_TRACE=client_channel,http2
# For comprehensive tracingexport GRPC_TRACE=api,call_error,channel,client_channel,connectivity_state,http,http2_stream,tcp,transport_security
# Run your script
python your_script.pyNote: These environment variables must be set before Python starts, as the gRPC C++ core reads them at module import time.
| Issue | Solution |
|---|---|
| Connection timeout | Check network connectivity, firewall rules, and ensure port 80/443 is open |
| Authentication failure | Verify username (email) and access token are correct |
| 503 Service Unavailable | Cluster may be suspended; enable auto_resume=True |
| 456 Strategy Error | Automatic blue-green failover will handle this |
| Memory issues with large results | Use fetchall_buffer() instead of fetchall() |
| gRPC message size errors | Configure grpc_options with appropriate message size limits |
| fastbinary import error | Install system dependencies (python3-devel, gcc-c++) or set require_fastbinary=False |
See TECH_DOC.md for detailed technical documentation.