Skip to content

Disruptor Queue Batching (DON'T MERGE YET) - #704

Closed
revans2 wants to merge 1 commit into
apache:masterfrom
revans2:disruptor-batching
Closed

Disruptor Queue Batching (DON'T MERGE YET)#704
revans2 wants to merge 1 commit into
apache:masterfrom
revans2:disruptor-batching

Conversation

@revans2

Copy link
Copy Markdown
Contributor

This is mostly a proof of concept, as a comparison to #694. They may complement each other. Of there may be a lot of overlap between them. I want to run some more tests. Currently I have a few benchmarks based on this branch showing the CPU savings, and increased performance by doing some batching.

The following were all using the modified version of wordcount on my branch. They were run on my Mac Book Pro with 8 cores. The first number is the max spout pending. The second number is the batch size, and the third number is the disruptor queue timeout.

Also before merging anything I want to explore reducing the number of threads and queues. Because if context switching is killing us having less threads should make it a lot better, and may reduce the need for this a lot.

wc-1-1-100 8 cores System - 30% User - 60% Idle - 10%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 41s649406494011.852642400
5m 7s14578014578010.2541448400
7m 4s2096802096809.9322086400

wc-1-100-100 8 cores System - 4% User - 5% Idle - 91%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
7m 0s154015401138.41217000

wc-10-1-100 8 cores System - 30% User - 65% Idle - 6%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 46s665406654099.480654000
5m 6s14230014230095.1011415400
7m 1s20666020666093.9512066000

wc-10-2-100 8 cores System - 30% User - 65% Idle - 6%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
3m 9s943409434090.511939400
5m 0s15542015542090.2811547000
7m 2s22302022302090.0472218000

wc-10-3-100 8 cores System - 30% User - 65% Idle - 6%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 42s705607056098.507701200
5m 2s15004015004092.6561500000
7m 3s21740021740090.7772189200

wc-200-1-100 8 cores System - 50% User - 50% Idle - 0%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 49s46440464403083.811451600
5m 6s94920949202990.687933400
7m 3s1339201339203008.9931326000

wc-200-100-100 8 cores System - 20% User - 30-40% Idle - 40-50%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 48s48680486802962.734474800
5m 3s99840998402885.372976400
7m 3s1429601429602845.1311408600

wc-400-1-100 8 cores System - 50% User - 50% Idle - 0%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
5m 6s1008801008805725.489978800
7m 13s1414001414005780.9841381600

wc-400-100-100 8 cores System - 25% User - 40-50% Idle - 20-30%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 46s63960639604206.964627000
5m 0s1367201367204011.0311354400
7m 1s1976401976403982.9461965200

wc-1000-200-100 8 cores System - 28% User - 68% Idle - 3%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 43s77000770008417.592734800
5m 4s1585601585608439.0141558000
7m 4s2302402302408456.8932265800

wc-1000-1-100 8 cores System - 50% User - 50% Idle - 0%

TIMEEMITTEDTRANSCOMP LATACKEDFAILED
2m 43s399803998018441.4553514020
5m 5s845208452016939.5967960020
7m 6s12172012172016697.25011662020

@revans2

Copy link
Copy Markdown
ContributorAuthor

I have done a lot of testing on this, and we need batching, but this is not the correct way to do it. I have a proof of concept with some micro benchmarks, but I am going to probably do a pull request for a fixed version shortly.

@revans2revans2 closed this Sep 18, 2015
@mjsax

Copy link
Copy Markdown
Member

"but this is not the correct way to do it" Do you refer to batching I work on or you own changes?

@revans2

Copy link
Copy Markdown
ContributorAuthor

@mjsax Sorry about the confusion. I have been playing around with batching and I was not referring to your code changes. Disruptor offers two different ways of batching. One is on the read side, which this patch is about. That did not work in improving the efficiency/throughput of the queue. Instead the batching has to happen on the enqueue side to reduce the load it is placing on the OS in signaling/waking up the worker thread.

I have some micro benchmarks that show this, and how much of a performance improvement we can potentially expect to see. I have been able to do over 2,000,000 word count sentences per second on my Mac Book Pro laptop, where as storm is much lower. I have not finished collecting numbers, and I am working on some code to do this in storm itself so we can see the impact it can have.

The main reason why I like this approach over batching at the spout is that the code is isolated to just the queue itself, and should not have an impact of the rest of storm, except potentially to remove some of the other code changes that were put into storm initially to try and improve the throughput.

But I really want to have a working prototype with performance numbers before I try to compare the two approaches.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@revans2@mjsax