Skip to content

Reduce the device bubble introduced by heavy loop synchronization in coalesced fetch/release(z3_leaf_module) - #6694

Merged
loadams merged 48 commits into
deepspeedai:masterfrom
inkcherry:reduce_coalesced_fetch_bubble
Jan 6, 2025
Merged

Reduce the device bubble introduced by heavy loop synchronization in coalesced fetch/release(z3_leaf_module)#6694
loadams merged 48 commits into
deepspeedai:masterfrom
inkcherry:reduce_coalesced_fetch_bubble

Conversation

@inkcherry

@inkcherry inkcherry commented Oct 31, 2024

Copy link
Copy Markdown
Contributor

depend on #6649

When performing fetch/release operations on Z3 leaf modules, the loop time is excessively long in fine-grained module. Compared to non-leaf modules, Z3 leaf modules may include a larger number of parameters. Although each loop unit does not consume much time, the overall loop length can be significant.
image
The fetch time is impacted by:

Post-allgather operations (narrow, slice ,cat, difficult to avoid)
Memory pressure(record_stream/fetch event create&sync)
The release time is impacted by:
slice
Free parameter record_stream

Considering the fine-grained leaf modules, where each parameter is relatively small, we can treat the parameters within each leaf module as a unified entity to handle memory pressure. This approach can approximately halve the CPU time required for fetch/release operations.

@inkcherry
inkcherry requested a review from loadams as a code owner November 12, 2024 08:16
@inkcherry

Copy link
Copy Markdown
Contributor Author

Hi , @tjruwase. Some models use a large number of MoE experts, but the total number of parameters for MoE remains unchanged. This PR handles such cases as a whole to optimize memory management and reduce bubbles caused by long loops. when zero_module_granularity_threshold is enabled.

Do you have any suggestions?

@tjruwase
tjruwase removed the request for review from awan-10 December 23, 2024 18:39
Comment thread deepspeed/runtime/zero/partition_parameters.py Outdated
Comment thread deepspeed/runtime/zero/parameter_offload.py Outdated
@inkcherry

Copy link
Copy Markdown
Contributor Author

hi, @tjruwase ,It seems that the OOM in the CI is not caused by this PR. Could you please help retrigger it? Thanks!

@loadams

loadams commented Jan 2, 2025

Copy link
Copy Markdown
Collaborator

hi, @tjruwase ,It seems that the OOM in the CI is not caused by this PR. Could you please help retrigger it? Thanks!

Hi @inkcherry - thanks, this is a known issue we are working on resolving with our CI nodes. I'll re-trigger the CI.

@loadams
loadams enabled auto-merge January 6, 2025 17:41
@loadams
loadams added this pull request to the merge queue Jan 6, 2025
Merged via the queue into deepspeedai:master with commit b0040b6 Jan 6, 2025
siqi654321 pushed a commit to siqi654321/DeepSpeed that referenced this pull request Feb 7, 2025
…coalesced fetch/release(z3_leaf_module) (deepspeedai#6694)

depend on deepspeedai#6649

When performing fetch/release operations on Z3 leaf modules, the loop
time is excessively long in fine-grained module. Compared to non-leaf
modules, Z3 leaf modules may include a larger number of parameters.
Although each loop unit does not consume much time, the overall loop
length can be significant.

![image](https://github.com/user-attachments/assets/9891835a-2620-47f3-aba6-ea22b8905d1c)
**The fetch time is impacted by:**

Post-allgather operations (narrow, slice ,cat, difficult to avoid)
Memory pressure(record_stream/fetch event create&sync)
**The release time is impacted by:**
slice
Free parameter record_stream

Considering the fine-grained leaf modules, where each parameter is
relatively small, we can treat the parameters within each leaf module as
a unified entity to handle memory pressure. This approach can
approximately halve the CPU time required for fetch/release operations.

---------

Co-authored-by: Ma, Guokai <guokai.ma@gmail.com>
Co-authored-by: Logan Adams <114770087+loadams@users.noreply.github.com>
Co-authored-by: Olatunji Ruwase <olruwase@microsoft.com>
Signed-off-by: siqi <siqi@tecorigin.com>
traincheck-team pushed a commit to traincheck-team/DeepSpeed that referenced this pull request Feb 9, 2025
…coalesced fetch/release(z3_leaf_module) (deepspeedai#6694)

depend on deepspeedai#6649

When performing fetch/release operations on Z3 leaf modules, the loop
time is excessively long in fine-grained module. Compared to non-leaf
modules, Z3 leaf modules may include a larger number of parameters.
Although each loop unit does not consume much time, the overall loop
length can be significant.

![image](https://github.com/user-attachments/assets/9891835a-2620-47f3-aba6-ea22b8905d1c)
**The fetch time is impacted by:**

Post-allgather operations (narrow, slice ,cat, difficult to avoid)
Memory pressure(record_stream/fetch event create&sync)
**The release time is impacted by:**
slice
Free parameter record_stream

Considering the fine-grained leaf modules, where each parameter is
relatively small, we can treat the parameters within each leaf module as
a unified entity to handle memory pressure. This approach can
approximately halve the CPU time required for fetch/release operations.

---------

Co-authored-by: Ma, Guokai <guokai.ma@gmail.com>
Co-authored-by: Logan Adams <114770087+loadams@users.noreply.github.com>
Co-authored-by: Olatunji Ruwase <olruwase@microsoft.com>
mauryaavinash95 pushed a commit to DataStates/DeepSpeed that referenced this pull request Mar 20, 2025
…coalesced fetch/release(z3_leaf_module) (deepspeedai#6694)

depend on deepspeedai#6649

When performing fetch/release operations on Z3 leaf modules, the loop
time is excessively long in fine-grained module. Compared to non-leaf
modules, Z3 leaf modules may include a larger number of parameters.
Although each loop unit does not consume much time, the overall loop
length can be significant.

![image](https://github.com/user-attachments/assets/9891835a-2620-47f3-aba6-ea22b8905d1c)
**The fetch time is impacted by:**

Post-allgather operations (narrow, slice ,cat, difficult to avoid)
Memory pressure(record_stream/fetch event create&sync)
**The release time is impacted by:**
slice
Free parameter record_stream

Considering the fine-grained leaf modules, where each parameter is
relatively small, we can treat the parameters within each leaf module as
a unified entity to handle memory pressure. This approach can
approximately halve the CPU time required for fetch/release operations.

---------

Co-authored-by: Ma, Guokai <guokai.ma@gmail.com>
Co-authored-by: Logan Adams <114770087+loadams@users.noreply.github.com>
Co-authored-by: Olatunji Ruwase <olruwase@microsoft.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants