IP Library › Granted Patent US 12,165,654
Granted Patent B2
US 12,165,654 · App. 17/595,472 · Granted Dec 10, 2024

Multi-speaker diarization of audio input using a neural network

Inventors: Yusuke Fujita (Tokyo, JP); Shinji Watanabe (Ellicott City, MD); Naoyuki Kanda (Tokyo, JP); Shota Horiguchi (Tokyo, JP)
Assignees: The Johns Hopkins University; Hitachi, Ltd.
G10L17/18G06N3/045G10L17/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,165,654
App. No.
17/595,472
Granted
Dec 10, 2024
Kind
B2
Abstract

An audio analysis platform may receive a portion of an audio input, wherein the audio input corresponds to audio associated with a plurality of speakers. The audio analysis platform may process, using a neural network, the portion of the audio input to determine voice activity of the plurality of speakers during the portion of the audio input, wherein the neural network is trained using reference audio data and reference diarization data corresponding to the reference audio data. The audio analysis platform may determine, based on the neural network being used to process the portion of the audio input, a diarization output associated with the portion of the audio input, wherein the diarization output indicates individual voice activity of the plurality of speakers. The audio analysis platform may provide the diarization output to indicate the individual voice activity of the plurality of speakers during the portion of the audio input.

Claims (86)

1. A method, comprising:

receiving, by a device, a portion of an audio input,

wherein the audio input corresponds to audio associated with a plurality of speakers;

processing, by the device and using a neural network, the portion of the audio input to determine voice activity of the plurality of speakers during the portion of the audio input,

wherein the neural network is trained using:

reference audio data, and

reference diarization data corresponding to the reference audio data,

wherein the reference diarization data indicates voice activity of individual reference speakers relative to timing of the reference audio data;

determining, by the device and based on the neural network being used to process the portion of the audio input, a diarization output associated with the portion of the audio input,

wherein processing the portion of the audio input comprises:

determining permutations of voice activity of the plurality of speakers; and

performing a cross entropy analysis of the permutations,

wherein the diarization output is determined to be a permutation of the permutations, of voice activity of the plurality of speakers, that satisfies a threshold permutation-invariant loss value according to the cross entropy analysis, and

wherein the diarization output indicates individual voice activity of the plurality of speakers; and

providing, by the device, the diarization output to indicate the individual voice activity of the plurality of speakers during the portion of the audio input.

2. The method of claim 1 , wherein processing the portion of the audio input comprises:

determining permutations of voice activity of the plurality of speakers; and

performing a loss analysis of the permutations of voice activity of the plurality of speakers,

wherein the diarization output is determined to be a permutation that satisfies a threshold permutation-invariant loss value according to the loss analysis.

3. The method of claim 1 , wherein the reference diarization data of the neural network is selected to train the neural network based on a permutation invariant analysis of a plurality of reference audio samples.

4. The method of claim 3 , wherein, when determining the diarization output, the permutation invariant analysis causes the neural network to reduce a permutation-invariant loss,

wherein the permutation-invariant loss is reduced by causing the neural network to select the diarization output from the permutations of voice activity of the plurality of speakers.

5. The method of claim 1 , wherein the portion of the audio input is a first portion of the audio input, the diarization output is a first diarization output, and the neural network is a first neural network, and

wherein the method further comprises:

receiving a second portion of the audio input;

processing, using a second neural network, the second portion of the audio input to determine voice activity of the plurality of speakers during the second portion of the audio input;

determining, based on the second neural network being used to process the second portion of the audio input, a second diarization output associated with the second portion of the audio input; and

providing the second diarization output to indicate the individual voice activity of the plurality of speakers during the second portion of the audio input.

6. The method of claim 5 , wherein the first neural network and the second neural network are separate instances of the same neural network.

7. A device, comprising:

one or more memories; and

one or more processors communicatively coupled to the one or more memories, configured to:

receive a portion of an audio input,

wherein the audio input corresponds to audio associated with a plurality of speakers;

process, using a neural network, the portion of the audio input to determine voice activity of the plurality of speakers during the portion of the audio input,

wherein the neural network is trained using:

reference audio data, and

reference diarization data corresponding to the reference audio data,

 wherein the reference diarization data indicates voice activity of individual reference speakers relative to timing of the reference audio data;

determine, based on the neural network being used to process the portion of the audio input, a diarization output associated with the portion of the audio input,

wherein processing the portion of the audio input comprises:

determining permutations of voice activity of the plurality of speakers; and

performing a cross entropy analysis of the permutations of voice activity of the plurality of speakers,

 wherein the diarization output is determined to be the permutation, of voice activity of the plurality of speakers, that satisfies a threshold permutation-invariant loss value according to the cross entropy analysis, and

wherein the diarization output indicates individual voice activity of the plurality of speakers; and

provide the diarization output to indicate the individual voice activity of the plurality of speakers during the portion of the audio input.

8. The device of claim 7 , wherein the reference audio data of the neural network is selected to train the neural network based on a permutation invariant analysis of a plurality of reference audio samples.

9. The device of claim 7 , wherein the diarization output indicates that two or more of the plurality of speakers are simultaneously actively speaking during the portion of the audio input.

10. The device of claim 7 , wherein the audio input comprises an audio stream, and

wherein the diarization output is provided in real-time relative to streaming playback of the portion of the audio stream.

11. The device of claim 7 , wherein the neural network comprises at least one of:

a recurrent neural network, or

a long short-term memory neural network.

12. The device of claim 7 , wherein the one or more processors are further configured to:

synchronize the diarization output with the audio input to generate an annotated audio output,

wherein the diarization output is provided within the annotated audio output.

13. The device of claim 12 , wherein the annotated audio output enables one or more of the plurality of speakers to be identified as an active speaker in association with the portion of the audio input.

14. The device of claim 8 , wherein, when determining the diarization output, the permutation invariant analysis causes the neural network to reduce a permutation-invariant loss,

wherein the permutation-invariant loss is reduced by causing the neural network to select the diarization output from the permutations of voice activity of the plurality of speakers.

15. A non-transitory computer-readable medium storing instructions, the instructions comprising:

one or more instructions that, when executed by one or more processors, cause the one or more processors to:

receive a portion of an audio input,

wherein the audio input corresponds to audio associated with a plurality of speakers;

process, using a neural network, the portion of the audio input to determine voice activity of the plurality of speakers during the portion of the audio input,

wherein the neural network is trained using:

reference audio data, and

reference diarization data corresponding to the reference audio data,

 wherein the reference diarization data indicates voice activity of individual reference speakers relative to timing of the reference audio data;

determine, based on the neural network being used to process the portion of the audio input, a diarization output associated with the portion of the audio input,

wherein processing the portion of the audio input comprises:

determining permutations of voice activity of the plurality of speakers; and

performing a cross entropy analysis of the permutations of voice activity of the plurality of speakers,

 wherein the diarization output is determined to be a permutation, of voice activity of the plurality of speakers, that satisfies a threshold permutation-invariant loss value according to the cross entropy analysis, and

wherein the diarization output indicates individual voice activity of the plurality of speakers; and

provide the diarization output to indicate the individual voice activity of the plurality of speakers during the portion of the audio input.

16. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the one or more processors to process the portion of the audio input, further cause the one or more processors to:

determine permutations of voice activity of the plurality of speakers; and

perform a loss analysis of the permutations of voice activity of the plurality of speakers,

wherein the diarization output is determined to be a permutation that satisfies a threshold permutation-invariant loss value according to the loss analysis.

17. The non-transitory computer-readable medium of claim 15 , wherein the reference diarization data of the neural network is selected to train the neural network based on a permutation invariant analysis of a plurality of reference audio samples.

18. The non-transitory computer-readable medium of claim 17 , wherein, when determining the diarization output, the permutation invariant analysis causes the neural network to reduce a permutation-invariant loss,

wherein the permutation-invariant loss is reduced by causing the neural network to select the diarization output from the permutations of voice activity of the plurality of speakers.

19. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions further cause the one or more processors to:

synchronize the diarization output with the audio input to generate an annotated audio output,

wherein the annotated audio output enables one or more of the plurality of speakers to be identified as an active speaker in association with the portion of the audio input.

20. The non-transitory computer-readable medium of claim 15 , wherein the reference audio data of the neural network is selected to train the neural network based on a permutation invariant analysis of a plurality of reference audio samples.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2022
From: FUJITA, YUSUKE; KANDA, NAOYUKI; HORIGUCHI, SHOTA
To: HITACHI, LTD.
Reel/Frame 059217/0521 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2022
From: WATANABE, SHINJI
To: THE JOHNS HOPKINS UNIVERSITY
Reel/Frame 058873/0119 →
Continuity (2)
Provisional Application 62896392 · Sep 5, 2019
Related Publication 20220254352A1 · Aug 11, 2022