🚀 Feature request
I would like to finetune t5-11b model on my dataset, but found that it doesn't fit in TPU or GPU memory - colab notebook just crash when I run it.
I tried to find a ready model parallelism solution. First I found this PR:
#3578
but it seems it haven't released. I tried to merge it to master branch locally, and use it, but it's crashed.
Also I have found Eisen library that propose "model parallelism with one code line", but works only for models with only one input ( t5 have 2 inputs - tokens and mask).
I need to distribute model on several GPU, and I see somebody tried to perform it. If this development (pull request 3578) is still in process, can you tell is there are any plans to release it?
🚀 Feature request
I would like to finetune t5-11b model on my dataset, but found that it doesn't fit in TPU or GPU memory - colab notebook just crash when I run it.
I tried to find a ready model parallelism solution. First I found this PR:
#3578
but it seems it haven't released. I tried to merge it to master branch locally, and use it, but it's crashed.
Also I have found Eisen library that propose "model parallelism with one code line", but works only for models with only one input ( t5 have 2 inputs - tokens and mask).
I need to distribute model on several GPU, and I see somebody tried to perform it. If this development (pull request 3578) is still in process, can you tell is there are any plans to release it?