IP Library Granted Patent US 11,443,748
Granted Patent B2
US 11,443,748 · App. 16/807,818 · Granted Sep 13, 2022

Metric learning of speaker diarization

Inventor: Masayuki Suzuki (Tokyo, JP)
Assignee: International Business Machines Corporation
G10L17/04G06N20/00G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,443,748
App. No.
16/807,818
Granted
Sep 13, 2022
Kind
B2
Abstract

A computer-implemented method includes obtaining, using a hardware processor, training data including utterances of speakers and performing tasks to train a machine learning model that converts an utterance into a feature vector, each task using one subset of multiple subsets of training data. The subsets of training data include a first subset of training data including utterances of a first number of speakers and at least one second subset of training data. Each second subset of training data includes utterances of a number of speakers that is less than the first number of speakers.

Claims (42)

1. A computer-implemented method comprising:

obtaining, using a hardware processor, training data stored on one or more computer readable storage mediums, the training data including a plurality of utterances of a plurality of speakers; and

performing, using the hardware processor, a plurality of tasks to train a machine learning model that converts an utterance of the plurality of utterances into a feature vector, each task using one of a plurality of subsets of training data, wherein the plurality of subsets of training data includes:

a first subset of training data including utterances of a first number of speakers among the plurality of speakers, and

at least one second subset of training data, each second subset including utterances of a number of speakers among the plurality of speakers that is less than the first number of speakers among the plurality of speakers,

wherein performing the plurality of tasks to train the machine learning model includes performing one task using the first subset of training data and two or more tasks using the at least one second subset of training data.

2. The computer-implemented method of claim 1 , wherein performing the plurality of tasks to train the machine learning model includes performing the plurality of tasks according to a multi-task training technique.

3. The computer-implemented method of claim 1 , wherein

the machine learning model includes a first model for converting an utterance of the plurality of utterances into a feature vector and a second model for identifying a speaker of the plurality of speakers from a feature vector, and

each utterance of the plurality of utterances in the training data is paired with an identification of a speaker of the plurality of speakers corresponding thereto.

4. The computer-implemented method of claim 3 , further comprising producing, using the hardware processor, a converter that converts an utterance of a speaker of the plurality of speakers into a feature vector by training the first model.

5. The computer-implemented method of claim 3 , wherein performing the plurality of tasks includes using a value of at least one hyperparameter of the task using the first subset of training data that is different from the value of the at least one hyperparameter of the task using the at least one second subset of training data.

6. The computer-implemented method of claim 5 , wherein the at least one hyperparameter is a margin of loss function.

7. The computer-implemented method of claim 1 , wherein the utterances of the second subset of training data are recorded in a substantially similar acoustic environment.

8. The computer-implemented method of claim 1 , wherein the utterances of the second subset of training data are obtained from a single continuous recording.

9. The computer-implemented method of claim 1 , wherein the utterances of the first subset of training data are obtained by combining two or more audio recordings.

10. A computer program product including one or more computer-readable storage mediums collectively storing program instructions that are executable by a processor or programmable circuitry to cause the processor or programmable circuitry to perform operations comprising:

obtaining training data including a plurality of utterances of a plurality of speakers; and

performing a plurality of tasks to train a machine learning model that converts an utterance of the plurality of utterances into a feature vector, each task using one of a plurality of subsets of training data, wherein the plurality of subsets of training data includes:

a first subset of training data including utterances of a first number of speakers among the plurality of speakers, and

at least one second subset of training data, each second subset including utterances of a number of speakers among the plurality of speakers that is less than the first number of speakers among the plurality of speakers,

wherein performing the plurality of tasks to train the machine learning model includes performing one task using the first subset of training data and two or more tasks using the at least one second subset of training data.

11. The computer program product of claim 10 , wherein

the machine learning model includes a first model for converting an utterance of the plurality of utterances into a feature vector and a second model for identifying a speaker of the plurality of speakers from a feature vector, and

each utterance of the plurality of utterances in the training data is paired with an identification of a speaker of the plurality of speakers corresponding thereto.

12. The computer program product of claim 11 , wherein the operations further comprise producing a converter that converts an utterance of a speaker of the plurality of speakers into a feature vector by training the first model.

13. The computer program product of claim 11 , wherein performing the plurality of tasks includes using a value of at least one hyperparameter of the task using the first subset of training data that is different from the value of the at least one hyperparameter of the task using the at least one second subset of training data.

14. The computer program product of claim 10 , wherein the utterances of the second subset of training data are obtained from a single continuous recording.

15. An apparatus comprising:

a processor or programmable circuitry; and

one or more computer readable storage mediums collectively including instructions that, when executed by the processor or the programmable circuitry, cause the processor or the programmable circuitry to:

obtain training data including a plurality of utterances of a plurality of speakers; and

perform a plurality of tasks to train a machine learning model that converts an utterance of the plurality of utterances into a feature vector, each task using one of a plurality of subsets of training data, wherein the plurality of subsets of training data includes:

a first subset of training data including utterances of a first number of speakers among the plurality of speakers, and

at least one second subset of training data, each second subset including utterances of a number of speakers among the plurality of speakers that is less than the first number of speakers among the plurality of speakers,

wherein performing the plurality of tasks to train the machine learning model includes performing one task using the first subset of training data and two or more tasks using the at least one second subset of training data.

16. The apparatus of claim 15 , wherein

the machine learning model includes a first model for converting an utterance of the plurality of utterances into a feature vector and a second model for identifying a speaker of the plurality of speakers from a feature vector, and

each utterance of the plurality of utterances in the training data is paired with an identification of a speaker of the plurality of speakers corresponding thereto.

17. The apparatus of claim 16 , wherein the instructions further cause the processor or the programmable circuitry to produce a converter that converts an utterance of a speaker of the plurality of speakers into a feature vector by training the first model.

18. The apparatus of claim 16 , wherein performing the plurality of tasks includes using a value of at least one hyperparameter of the task using the first subset of training data that is different from the value of the at least one hyperparameter of the task using the at least one second subset of training data.

19. The apparatus of claim 15 , wherein the utterances of the second subset of training data are obtained from a single continuous recording.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2020
From: SUZUKI, MASAYUKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051996/0922 →
Continuity (1)
Related Publication 20210280196A1 · Sep 9, 2021