Problem with Utterance-level Prosody extractor of DelightfulTTS #7

vietvq-vbee · 2022-01-31T05:24:26Z

I've recently been experimenting with your implementation of DelightfulTTS and the voice quality is awesome. However I found out that the embedding vector output of Utterance-level Prosody extractor is very small, making the that of Utterance-level Prosody predictor small as well (L2 is roughly 12 and each element in the vector is roughly 0.2 to 0.3). Vectors with element close to zero means this layer mostly doesn't add any information at all. Have you find any solution to this?

keonlee9420 · 2022-02-18T14:08:01Z

Hi @vietvq-vbee , thanks for sharing and sorry for late response. I just updated repo (v0.2.0) with some improvements, but still prosody modelings including DelightfulTTS are not yet resolved (WIP). I'll take a look with your insight and update the repo if I can make it work!

vietvq-vbee · 2022-03-24T14:40:38Z

@keonlee9420 I think I've found the source of the problem mentioned above. My colleague and I suspect this is because the Conformer layers use ReLU as activation function ([0, inf]) and UtteranceLevelProsodyEncoder uses tanh as activation function ([-1, 1]), meaning the maximum value of UtteranceLevelProsodyEncoder is still very small comparing to the average value of Conformer layers.

We haven't conducted experiments where we replace tanh -> ReLU or LeakyReLU since currently we're concatenating them, but I'll inform you via this discussion ASAP :)

keonlee9420 · 2022-03-26T01:04:40Z

Great! Looking forward to seeing how the results turn out :)

devangsrammohan · 2023-10-25T20:30:27Z

I know I'm a bit late to the party, but I'm curious to know if either of you were able to resolve the issues with utterance level prosody extraction? When I train a model, it appears to ignore the utterance prosody embedding altogether.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Problem with Utterance-level Prosody extractor of DelightfulTTS #7

Problem with Utterance-level Prosody extractor of DelightfulTTS #7

vietvq-vbee commented Jan 31, 2022 •

edited

Loading

keonlee9420 commented Feb 18, 2022

vietvq-vbee commented Mar 24, 2022

keonlee9420 commented Mar 26, 2022

devangsrammohan commented Oct 25, 2023

Problem with Utterance-level Prosody extractor of DelightfulTTS #7

Problem with Utterance-level Prosody extractor of DelightfulTTS #7

Comments

vietvq-vbee commented Jan 31, 2022 • edited Loading

keonlee9420 commented Feb 18, 2022

vietvq-vbee commented Mar 24, 2022

keonlee9420 commented Mar 26, 2022

devangsrammohan commented Oct 25, 2023

vietvq-vbee commented Jan 31, 2022 •

edited

Loading