Implement xm.rendezvous with XLA collective communication - #4181
Merged
Conversation
ronghanghu
reviewed
Nov 10, 2022
will-cromar
force-pushed
the
wcromar/xla-rendezvous
branch
from
November 10, 2022 20:53
78ecc43 to
585982d
Compare
xm.rendezvous with XLA collective communicationxm.rendezvous with XLA collective communication
will-cromar
marked this pull request as ready for review
November 10, 2022 21:00
Contributor
|
I tested this on v4-8 and v4-4096 and it seems to work successfully on both accelerator types. Using |
JackCaoG
reviewed
Nov 11, 2022
JackCaoG
approved these changes
Nov 15, 2022
JackCaoG
left a comment
Collaborator
There was a problem hiding this comment.
Anything in the PJRT README we should update?
Collaborator
Author
|
I'll update the readme after #4193 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
We have found that
gloodoesn't scale effectively to large pod sizes, and it's not easily possible to usetorch.distributedin a multithreaded context such as TPU v3.xm.rendezvouswill now callxm.mark_stepto sync results from XLA.xm.rendezvous. Also remove initialization based onXRT_MESH_SERVICE_ADDRESSsince host 0 is not predictable anyway.xm.rendezvousfrom all replicas per XLA requirements:Computing the result of AllReduce requires having one input from each replica, so if one replica executes a AllReduce node more times than another, then the former replica will wait forever.This covers the vast majority of the usage ofrendezvousin our experience.Tested manually on a TPU v4-8 with 1 process and 4 threads to simulate a v3.