IP Library Granted Patent US 11,551,668
Granted Patent B1
US 11,551,668 · App. 17/138,362 · Granted Jan 10, 2023

Generating representations of speech signals using self-supervised learning

Inventors: Alexei Baevski (Redwood City, CA); Yuhao Zhou (Menlo Park, CA); Abdelrahman Mohamed (Redmond, WA); Michael Auli (Portola Valley, CA); Ronan Stéfan Collobert (Mountain View, CA); Alexis Conneau (Foster City, CA)
Assignee: Meta Platforms, Inc.
G10L15/063G06K9/6259G10L15/16G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,668
App. No.
17/138,362
Granted
Jan 10, 2023
Kind
B1
Abstract

In one embodiment, a method includes generating audio segments from a speech signal, generating latent representations that respectively correspond to the audio segments, the latent representations comprising a first subset and a second subset, generating quantized representations that respectively correspond to the latent representations, masking the second subset of the latent representations, using a machine-learning model to process the first subset of the latent representations and the masked second subset of the latent representations to generate contextualized representations that respectively correspond to the latent representations, pre-training the machine-learning model based on comparisons between (1) a subset of the contextualized representations that respectively correspond to the masked second subset of the latent representations and (2) a subset of the quantized representations that respectively correspond to the masked second subset of the latent representations, and training the pre-trained machine-learning model to perform a speech analysis task.

Claims (53)

1. A method comprising, by a computing system:

generating audio segments from a speech signal;

generating latent representations that respectively correspond to the audio segments, the latent representations comprising a first subset and a second subset;

generating quantized representations that respectively correspond to the latent representations;

masking the second subset of the latent representations;

using a machine-learning model to process the first subset of the latent representations and the masked second subset of the latent representations to generate contextualized representations that respectively correspond to the latent representations;

pre-training the machine-learning model based on comparisons between (1) a subset of the contextualized representations that respectively correspond to the masked second subset of the latent representations and (2) a subset of the quantized representations that respectively correspond to the masked second subset of the latent representations; and

training the pre-trained machine-learning model to perform a speech analysis task.

2. The method of claim 1 , further comprising:

normalizing the speech signal to zero mean and unit variance.

3. The method of claim 1 , wherein generating the audio segments is based on one or more time-steps, and wherein each of the one or more time-steps comprises an amount of time.

4. The method of claim 1 , wherein generating the latent representations is based on a multi-layer convolutional neural network.

5. The method of claim 1 , wherein generating the quantized representations is based on product quantization.

6. The method of claim 1 , wherein generating each of the quantized representations comprises:

accessing a plurality of codebooks, wherein each of the plurality of codebooks comprises a plurality of vector entries;

selecting one vector entry from each of the plurality of codebooks;

concatenating the plurality of vector entries to generate a concatenated vector; and

applying a linear transformation to the concatenated vector to generate the quantized representation.

7. The method of claim 6 , wherein generating each of the quantized representations is based on a diversity loss function, and wherein the diversity loss function optimizes a probability of selecting each of the plurality of vector entries in each of the plurality of codebooks to be equal.

8. The method of claim 1 , wherein each of the quantized representations is associated with a true quantized representation and one or more distractors, wherein generating the contextualized representations is based on a contrastive loss function, and wherein the contrastive loss function optimizes a contextualized representation to be similar to a corresponding true quantized representation but different from the one or more associated distractors.

9. The method of claim 1 , wherein pre-training the machine-learning model is on a plurality of unlabeled training data.

10. The method of claim 1 , wherein training the pre-trained machine-learning model is based on one or more labeled training data, wherein the one or more labeled training data are associated with the speech analysis task.

11. The method of claim 1 , wherein the speech signal is based on a plurality of languages.

12. The method of claim 11 , wherein each of the latent representations is common to the plurality of languages.

13. The method of claim 11 , wherein each of the quantized representations is common to the plurality of languages.

14. The method of claim 11 , wherein each of the contextualized representations is common to the plurality of languages.

15. The method of claim 11 , wherein pretraining the machine-learning model is based on a plurality of unlabeled training data associated with the plurality of languages.

16. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

generate audio segments from a speech signal;

generate latent representations that respectively correspond to the audio segments, the latent representations comprising a first subset and a second subset;

generate quantized representations that respectively correspond to the latent representations;

mask the second subset of the latent representations;

use a machine-learning model to process the first subset of the latent representations and the masked second subset of the latent representations to generate contextualized representations that respectively correspond to the latent representations;

pre-train the machine-learning model based on comparisons between (1) a subset of the contextualized representations that respectively correspond to the masked second subset of the latent representations and (2) a subset of the quantized representations that respectively correspond to the masked second subset of the latent representations; and

train the pre-trained machine-learning model to perform a speech analysis task.

17. The media of claim 16 , wherein generating each of the quantized representations comprises:

accessing a plurality of codebooks, wherein each of the plurality of codebooks comprises a plurality of vector entries;

selecting one vector entry from each of the plurality of codebooks;

concatenating the plurality of vector entries to generate a concatenated vector; and

applying a linear transformation to the concatenated vector to generate the quantized representation.

18. A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

generate audio segments from a speech signal;

generate latent representations that respectively correspond to the audio segments, the latent representations comprising a first subset and a second subset;

generate quantized representations that respectively correspond to the latent representations;

mask the second subset of the latent representations;

use a machine-learning model to process the first subset of the latent representations and the masked second subset of the latent representations to generate contextualized representations that respectively correspond to the latent representations;

pre-train the machine-learning model based on comparisons between (1) a subset of the contextualized representations that respectively correspond to the masked second subset of the latent representations and (2) a subset of the quantized representations that respectively correspond to the masked second subset of the latent representations; and

train the pre-trained machine-learning model to perform a speech analysis task.

19. The system of claim 18 , wherein generating each of the quantized representations comprises:

accessing a plurality of codebooks, wherein each of the plurality of codebooks comprises a plurality of vector entries;

selecting one vector entry from each of the plurality of codebooks;

concatenating the plurality of vector entries to generate a concatenated vector; and

applying a linear transformation to the concatenated vector to generate the quantized representation.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2021
From: BAEVSKI, ALEXEI; ZHOU, YUHAO; MOHAMED, ABDELRAHMAN; AULI, MICHAEL; COLLOBERT, RONAN STEFAN; CONNEAU, ALEXIS
To: FACEBOOK, INC.
Reel/Frame 055006/0021 →