Scalable streaming analytics pipeline implementing Apriori, PCY, and SON algorithms on Amazon product data using Apache Kafka, DASK, and MongoDB. Achieves 75% faster preprocessing with real-time frequent itemset discovery and association rule mining.
pythonmachine-learningdata-miningkafkabig-datamongodbnosqldata-engineeringdaskassociation-rulesbatch-processingstreaming-datareal-time-processingapriori-algorithmfrequent-itemsetspcy-algorithmson-algorithm
-
Updated
May 6, 2026 - Jupyter Notebook