Skip to content

fix opt-350m shard loading issue in AutoTP - #3600

Merged
jeffra merged 7 commits into
deepspeedai:masterfrom
sywangyi:fix-opt-350m
Jul 27, 2023
Merged

jeffra merged 7 commits into
deepspeedai:masterfrom
sywangyi:fix-opt-350m

Conversation

@sywangyi

Copy link
Copy Markdown
Contributor

No description provided.

Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>
@sywangyi

Copy link
Copy Markdown
Contributor Author

@delock @tjruwase please help review

@delock

delock commented Jun 15, 2023

Copy link
Copy Markdown
Collaborator

@tjruwase @jeffra could assign a reviewer for this PR? This PR fix OPT checkpoint sharded loading with AutoTP and improve OPT+AutoTP usability, it is needed when run OPT models on CPU server with small memory.

@delock

delock commented Jul 6, 2023

Copy link
Copy Markdown
Collaborator

@RezaYazdaniAminabadi can you review this PR? This PR fix OPT sharded loading for AutoTP. Previously only OPT-125m has sharded checkpoint loading, with this fix OPT >350m will have sharded checkpoint loading as well.

@delock

delock commented Jul 18, 2023

Copy link
Copy Markdown
Collaborator

@RezaYazdaniAminabadi Hi, a quick check whether this PR is still under consideration. We have verified this PR for CPU accelerator and like to know whether it could be merged into master branch, thanks!

@molly-smith
molly-smith self-requested a review July 27, 2023 20:11
@jeffra
jeffra enabled auto-merge July 27, 2023 20:15
@jeffra
jeffra added this pull request to the merge queue Jul 27, 2023
Merged via the queue into deepspeedai:master with commit 76953a3 Jul 27, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants