Conversation
EMTTS disables the alignment prior by passing prior_scaling_factor=0.0 to the data loader. However, both data loader flavors rely on the prior being `None` in order to skip creation of the prior. Both solutions are correct, but as we work on optimizing the data loader to enable faster training, it's a good idea to skip the creation entirely. This change sets `prior_scaling_factor` to `None` in all cases for EMTTS. Signed-off-by: Fejgin, Roy <rfejgin@nvidia.com>
rfejgin
requested review from
Edresson,
blisc and
paarthneekhara
and removed request for
Edresson and
blisc
September 28, 2026 23:19
Collaborator
Author
|
/ok to test 7b8a01f |
Contributor
|
[🤖]: Hi @rfejgin 👋, We wanted to let you know that a CICD pipeline for this PR just finished successfully. So it might be time to merge this PR or get some approvals. |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
EMTTS disables the alignment prior by passing
prior_scaling_factor=0.0to the data loader. However, the data loader(s) rely on the prior beingNonein order to actually skip the creation of the prior tensor. The existing implementation does not have a correctness issue (as scaling the prior to zero has the same effect as not using one), but it does result in unnecessary compute and memory use.This PR sets
prior_scaling_factortoNonein all cases for EMTTS.Background: EMTTS training currently runs out of host memory if more than a few data workers (per task) are created. More data loaders may be needed for training speed, and we have observed GPUs being stalled, possibly due to insufficient workers. Hence it makes sense to pay attention to optimizing the data loader, which is how this was found.