Hi. I'm running into memory issues when trying to fine tune the CTRL language model on a 16 GB GPU. Is there any built-in support for splitting models across more than one GPU? I suppose that mapping the input embedding layer to a different GPU or even the CPU would do the trick.
Hi. I'm running into memory issues when trying to fine tune the CTRL language model on a 16 GB GPU. Is there any built-in support for splitting models across more than one GPU? I suppose that mapping the input embedding layer to a different GPU or even the CPU would do the trick.