Skip to content
This repository was archived by the owner on Jun 11, 2026. It is now read-only.

Detokenization parallelization public repo - #69

Merged
NickNickGo merged 6 commits into
microsoft:mainfrom
NickNickGo:detokenization_parallelization_public_repo
Dec 15, 2020
Merged

Detokenization parallelization public repo#69
NickNickGo merged 6 commits into
microsoft:mainfrom
NickNickGo:detokenization_parallelization_public_repo

Conversation

@NickNickGo

Copy link
Copy Markdown
Contributor

Moving #37 here.

Comment threadfastseq_cli/transformers_generate.py Outdated
Comment threadfastseq_cli/transformers_generate.py Outdated
@NickNickGo

Copy link
Copy Markdown
ContributorAuthor
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval321024NANA34.89|14.96|25.30NANA1238.3NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval641024NANA34.92|14.95|25.25NANA8711.8NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval1281024NANA34.96|14.98|25.28NANA8212.5NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval321024NANA34.90|14.95|25.30NANA1218.5NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval641024NANA34.93|14.95|25.26NANA8711.8NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval1281024NANA34.97|14.96|25.27NANA8112.6NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval321024NANA34.91|14.94|25.25NANA1228.4NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval641024NANA34.93|14.97|25.25NANA8711.8NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm.1k/rawval1281024NANA34.98|14.96|25.26NANA8112.6NA
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4facebook/mbart-large-en-rowmt_en_ro/rawval641984NA27.89NA|NA|NANANA2747.2NA
transformers_v3.0.2+fastseq_v0.0.4facebook/mbart-large-en-rowmt_en_ro/rawval641984NA27.89NA|NA|NANANA2517.9NA
transformers_v3.0.2+fastseq_v0.0.4facebook/mbart-large-en-rowmt_en_ro/rawval641984NA27.89NA|NA|NANANA2547.8NA
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval641984NA27.44NA|NA|NANANA16211.2NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval1281984NA27.38NA|NA|NANANA10717.1NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval641984NA27.44NA|NA|NANANA13413.8NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval1281984NA27.38NA|NA|NANANA10617.3NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval641984NA27.44NA|NA|NANANA12913.8NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval1281984NA27.38NA|NA|NANANA10717.1NA
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm.1k/rawval641024NANA35.18|15.03|25.03NANA6815.1NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm.1k/rawval1281024NANA35.20|15.14|25.07NANA6415.9NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm.1k/rawval641024NANA35.21|15.07|25.00NANA6815.1NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm.1k/rawval1281024NANA35.19|15.10|25.08NANA6515.8NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm.1k/rawval641024NANA35.21|15.04|25.04NANA6715.3NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm.1k/rawval1281024NANA35.18|15.10|25.06NANA6515.8NA
~

@JiushengChen

Copy link
Copy Markdown
Contributor

CNN 1k data set is too small now, result is not reliable. Please use full valid set.

@NickNickGo

Copy link
Copy Markdown
ContributorAuthor
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm/rawval3213344NANA44.80|21.64|31.17NANA17637.6NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm/rawval6413312NANA44.79|21.66|31.18NANA117411.3NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm/rawval12813312NANA44.78|21.64|31.16NANA107512.4NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm/rawval6413312NANA45.06|21.81|30.91NANA81216.4NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm/rawval12813312NANA45.05|21.79|30.90NANA72518.4NA

@feihugis

Copy link
Copy Markdown
Contributor

For bart-large-cnn, why are the numbers of input examples different for different batch_sizes? Could you also paste the result for the baseline?

@NickNickGo

Copy link
Copy Markdown
ContributorAuthor

For bart-large-cnn, why are the numbers of input examples different for different batch_sizes? Could you also paste the result for the baseline?

This is because pytorch dataloader API only supports number of samples to be multiple of batch size. This is why last batch is dropped.
https://pytorch.org/docs/stable/data.html

@NickNickGo

Copy link
Copy Markdown
ContributorAuthor

Jiusheng Chen (@JiushengChen) For Mbart and T5 , do we have larger dataset? I couldn't find it in benchmark scripts.

@feihugis

Copy link
Copy Markdown
Contributor

For bart-large-cnn, why are the numbers of input examples different for different batch_sizes? Could you also paste the result for the baseline?

This is because pytorch dataloader API only supports number of samples to be multiple of batch size. This is why last batch is dropped.
https://pytorch.org/docs/stable/data.html

The root cause may be here. Try to change it to be drop_last=False?

@JiushengChen

Copy link
Copy Markdown
Contributor

Jiusheng Chen (@JiushengChen) For Mbart and T5 , do we have larger dataset? I couldn't find it in benchmark scripts.

Yes, I have larger data in my local. Please leave out these two, I will update them today.
BTW, looks CI test failed, please take a look.

@NickNickGo

NickNickGo commented Dec 10, 2020

Copy link
Copy Markdown
ContributorAuthor

For bart-large-cnn, why are the numbers of input examples different for different batch_sizes? Could you also paste the result for the baseline?

This is because pytorch dataloader API only supports number of samples to be multiple of batch size. This is why last batch is dropped.
https://pytorch.org/docs/stable/data.html

The root cause may be here. Try to change it to be drop_last=False?

I already did.

@feihugis

Copy link
Copy Markdown
Contributor

For bart-large-cnn, why are the numbers of input examples different for different batch_sizes? Could you also paste the result for the baseline?

This is because pytorch dataloader API only supports number of samples to be multiple of batch size. This is why last batch is dropped.
https://pytorch.org/docs/stable/data.html

The root cause may be here. Try to change it to be drop_last=False?

I already did.

Could you please explain more? I saw your code here use drop_last=True. I guess that's why the last batch was dropped. Do you mean you have tried drop_last=False but the last batch was still dropped?

Comment threadfastseq/optimizer/fairseq/generate.py Outdated
Comment threadbenchmarks/models/hf_bart.sh Outdated
Comment threadbenchmarks/models/hf_distibart.sh Outdated
Comment on lines +23 to +25
grep -E "transformers_v3.0.2\+fastseq_v.* hf.sshleifer.distilbart-cnn-12-6.tar.gz cnn_dm.1k/raw val 64 " perf | awk '{s+=$13}END{print s/NR}' | bash range.sh 13 100
grep -E "transformers_v3.0.2\+fastseq_v.* hf.sshleifer.distilbart-cnn-12-6.tar.gz cnn_dm.1k/raw val 64 " perf | awk '{s+=$13}END{print s/NR}' | bash range.sh 15.2 100
# todo: bigger bs doesn't increase speed
grep -E "transformers_v3.0.2\+fastseq_v.* hf.sshleifer.distilbart-cnn-12-6.tar.gz cnn_dm.1k/raw val 128 " perf | awk '{s+=$13}END{print s/NR}' | bash range.sh 13.5 100
grep -E "transformers_v3.0.2\+fastseq_v.* hf.sshleifer.distilbart-cnn-12-6.tar.gz cnn_dm.1k/raw val 128 " perf | awk '{s+=$13}END{print s/NR}' | bash range.sh 15.9 100

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same here

Comment threadfastseq_cli/transformers_generate.py Outdated
Comment threadbenchmarks/models/hf_bart.sh Outdated
Comment threadbenchmarks/models/hf_distibart.sh Outdated
Comment threadbenchmarks/models/hf_distibart.sh Outdated
Comment threadREADME.md Outdated
Comment threadbenchmarks/models/hf_distibart.sh Outdated
@NickNickGo

Copy link
Copy Markdown
ContributorAuthor

Benchmarks on Larger dataset:

UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm/rawval3213368NANA44.80|21.65|31.19NANA18897.1NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm/rawval6413368NANA44.80|21.66|31.19NANA118811.3NA
transformers_v3.0.2+fastseq_v0.0.4facebook/bart-large-cnncnn_dm/rawval12813368NANA44.78|21.64|31.18NANA108212.4NA
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm/rawval6413368NANA45.07|21.81|30.91NANA81016.5NA
transformers_v3.0.2+fastseq_v0.0.4hf.sshleifer.distilbart-cnn-12-6.tar.gzcnn_dm/rawval12813368NANA45.05|21.80|30.90NANA72918.3NA
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4facebook/mbart-large-en-rowmt_en_ro/rawval648191NA56.19NA|NA|NANANA8979.1NA
transformers_v3.0.2+fastseq_v0.0.4facebook/mbart-large-en-rowmt_en_ro/rawval648191NA56.19NA|NA|NANANA8849.3NA
transformers_v3.0.2+fastseq_v0.0.4facebook/mbart-large-en-rowmt_en_ro/rawval648191NA56.19NA|NA|NANANA8859.3NA
UtilModelTaskSplitBatchSizeSamplesTokensBleuRougeLossPerplexityRuntime(seconds)Throughput(samples/s)Throughput(tokens/s)
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval648191NA56.93NA|NA|NANANA42519.3NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval1288191NA56.92NA|NA|NANANA35023.4NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval648191NA56.93NA|NA|NANANA43618.8NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval1288191NA56.92NA|NA|NANANA36222.6NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval648191NA56.93NA|NA|NANANA44218.5NA
transformers_v3.0.2+fastseq_v0.0.4t5-basewmt_en_ro/rawval1288191NA56.92NA|NA|NANANA34024.1NA

@NickNickGo

NickNickGo commented Dec 11, 2020

Copy link
Copy Markdown
ContributorAuthor

Before/After

...
DistilBart13.818.3
T513.823.4
BART11.412.4
Mbart8.99.3

@NickNickGo

NickNickGo commented Dec 15, 2020

Copy link
Copy Markdown
ContributorAuthor

For bart-large-cnn, why are the numbers of input examples different for different batch_sizes? Could you also paste the result for the baseline?

This is because pytorch dataloader API only supports number of samples to be multiple of batch size. This is why last batch is dropped.
https://pytorch.org/docs/stable/data.html

The root cause may be here. Try to change it to be drop_last=False?

I already did.

Could you please explain more? I saw your code here use drop_last=True. I guess that's why the last batch was dropped. Do you mean you have tried drop_last=False but the last batch was still dropped?

Synced offline. Included last batch.

@NickNickGo
NickNickGo merged commit 8d217ee into microsoft:mainDec 15, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NickNickGo@JiushengChen@feihugis