IP Library › Granted Patent US 12,217,761
Granted Patent B2
US 12,217,761 · App. 17/515,480 · Granted Feb 4, 2025

Target speaker mode

Inventors: Yuhui Chen (San Jose, CA); Qiyong Liu (Singapore, SG); Zhengwei Wei (Jiangxi, CN); Yangbin Zeng (Zhejiang, CN)
Assignee: Zoom Video Communications, Inc.
G10L17/08G10L17/04G10L25/21G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,761
App. No.
17/515,480
Granted
Feb 4, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for target speaker extraction. A target speaker extraction system receives an audio frame of an audio signal. A multi-speaker detection model analyzes the audio frame to determine whether the audio frame includes only a single-speaker or multiple speakers. When the audio frame includes only a single-speaker, the system inputs the audio frame to a target speaker VAD model to suppress speech in the audio frame from a non-target speaker based on comparing the audio frame to a voiceprint of a target speaker. When the audio frame includes multiple speakers, the system inputs the audio frame to a speech separation model to separate the voice of the target speaker from a voice mixture in the audio frame.

Claims (82)

1. A computer-implemented method for target speaker extraction, comprising:

receiving, by a target speaker extraction system, an audio frame of an audio signal and a corresponding video, wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning (“ML”) model, a trained lip-movement-based (“LM-based”) target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model;

responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a single speaker within the audio frame:

inputting, by the target speaker extraction system, the audio frame and the video to the trained LM-based target speaker VAD ML model; and

suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, speech in the audio frame from a non-target speaker, wherein suppressing the speech in the audio from a non-target speaker comprises comparing the audio frame to a voiceprint of a target speaker; and

responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a plurality of speakers within the audio frame:

inputting, by the target speaker extraction system, the audio frame to the trained speech separation ML model; and

separating, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker from a voice mixture in the audio frame.

2. The method of claim 1 , wherein suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, speech in the audio frame from the non-target speaker further comprises:

determining, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a suppression ratio; and

suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the speech in the audio frame from the non-target speaker based on the suppression ratio.

3. The method of claim 2 , further comprising:

generating, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a voiceprint of the non-target speaker;

comparing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the voiceprint of the non-target speaker to the voiceprint of the target speaker to determine a similarity score; and

determining, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the suppression ratio based on the similarity score.

4. The method of claim 1 , wherein:

the method further comprises:

receiving, by the target speaker extraction system, a voice recording from a target speaker prior to a video conference; and

generating, by the target speaker extraction system, the voiceprint of the target speaker from the voice recording prior to the video conference; and

separating, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker from the voice mixture in the audio frame further comprises:

extracting, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker during the video conference based on the voiceprint of the target speaker.

5. The method of claim 1 , further comprising:

determining, by the target speaker extraction system, an energy of the audio signal, wherein the audio signal is received during a video conference;

determining, by the target speaker extraction system, speech by the target speaker within the audio frame based on the energy of the audio signal; and

generating, by the target speaker extraction system, a voiceprint of the target speaker from the audio signal.

6. The method of claim 5 , wherein determining, by the target speaker extraction system, the speech by the target speaker within the audio frame based on the energy of the audio signal comprises determining that the energy exceeds a threshold.

7. The method of claim 1 , wherein the target speaker extraction system comprises a trained voiceprint extraction ML model, and the method further comprises:

generating, by the trained voiceprint extraction ML model of the target speaker extraction system, a voiceprint of the target speaker; and

sharing, by the trained voiceprint extraction ML model, one or more weights associated with the voiceprint with the trained speech separation ML model.

8. A non-transitory computer readable medium comprising processor-executable instructions configured to cause one or more processors to:

receive, by a target speaker extraction system, an audio frame of an audio signal and a corresponding video, wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning (“ML”) model, a trained lip-movement-based (“LM-based”) target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model;

responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a single speaker within the audio frame:

input, by the target speaker extraction system, the audio frame and the video to the trained LM-based target speaker VAD ML model; and

suppress, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, speech in the audio frame from a non-target speaker, wherein suppressing the speech in the audio from a non-target speaker comprises comparing the audio frame to a voiceprint of a target speaker; and

responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a plurality of speakers within the audio frame:

input, by the target speaker extraction system, the audio frame to the trained speech separation ML model; and

separate, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker from a voice mixture in the audio frame.

9. The non-transitory computer readable medium of claim 8 , further comprising processor-executable instructions configured to cause the one or more processors to:

determine, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a suppression ratio; and

suppress, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the speech in the audio frame from the non-target speaker based on the suppression ratio.

10. The non-transitory computer readable medium of claim 9 , further comprising processor-executable instructions configured to cause the one or more processors to:

generate, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a voiceprint of the non-target speaker;

compare, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the voiceprint of the non-target speaker to the voiceprint of the target speaker to determine a similarity score; and

determine, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the suppression ratio based on the similarity score.

11. The non-transitory computer readable medium of claim 8 , further comprising processor-executable instructions configured to cause the one or more processors to:

receive, by the target speaker extraction system, a voice recording from a target speaker prior to a video conference; and

generate, by the target speaker extraction system, the voiceprint of the target speaker from the voice recording prior to the video conference; and

extract, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker during the video conference based on the voiceprint of the target speaker.

12. The non-transitory computer readable medium of claim 8 , further comprising processor-executable instructions configured to cause the one or more processors to:

determining, by the target speaker extraction system, an energy of the audio signal, wherein the audio signal is received during a video conference;

determining, by the target speaker extraction system, speech by the target speaker within the audio frame based on the energy of the audio signal; and

generating, by the target speaker extraction system, a voiceprint of the target speaker from the audio signal.

13. The non-transitory computer readable medium of claim 12 , wherein the processor-executable instructions configured to cause the one or more processors to determine, by the target speaker extraction system, the speech by the target speaker within the audio frame based on the energy of the audio signal further comprise further comprise processor-executable instructions configured to cause the one or more processors to determine that the energy exceeds a threshold.

14. The non-transitory computer readable medium of claim 8 , wherein the target speaker extraction system comprises a trained voiceprint extraction ML model, and further comprising processor-executable instructions configured to cause the one or more processors to:

generate, by the trained voiceprint extraction ML model of the target speaker extraction system, a voiceprint of the target speaker; and

share, by the trained voiceprint extraction ML model, one or more weights associated with the voiceprint with the trained speech separation ML model.

15. A target speaker extraction system comprising:

a non-transitory computer-readable medium; and

one or more processors communicatively coupled to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

receive, by a target speaker extraction system, an audio frame of an audio signal and a corresponding video, wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning (“ML”) model, a trained lip-movement-based (“LM-based”) target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model;

responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a single speaker within the audio frame:

input, by the target speaker extraction system, the audio frame and the video to the trained LM-based target speaker VAD ML model; and

suppress, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, speech in the audio frame from a non-target speaker, wherein suppressing the speech in the audio from a non-target speaker comprises comparing the audio frame to a voiceprint of a target speaker; and

responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a plurality of speakers within the audio frame:

input, by the target speaker extraction system, the audio frame to the trained speech separation ML model; and

separate, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker from a voice mixture in the audio frame.

16. The system of claim 15 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to-further configured to

determine, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a suppression ratio; and

suppress, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the speech in the audio frame from the non-target speaker based on the suppression ratio.

17. The system of claim 16 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

generate, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a voiceprint of the non-target speaker;

compare, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the voiceprint of the non-target speaker to the voiceprint of the target speaker to determine a similarity score; and

determine, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the suppression ratio based on the similarity score.

18. The system of claim 15 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

receive, by the target speaker extraction system, a voice recording from a target speaker prior to a video conference; and

generate, by the target speaker extraction system, the voiceprint of the target speaker from the voice recording prior to the video conference; and

extract, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker during the video conference based on the voiceprint of the target speaker.

19. The system of claim 15 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to further configured to:

determine, by the target speaker extraction system, an energy of the audio signal, wherein the audio signal is received during a video conference;

determine, by the target speaker extraction system, speech by the target speaker within the audio frame based on the energy of the audio signal; and

generate, by the target speaker extraction system, a voiceprint of the target speaker from the audio signal.

20. The system of claim 19 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to determine, by the target speaker extraction system, the speech by the target speaker within the audio frame based on the energy of the audio signal exceeding a threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2021
From: CHEN, YUHUI; LIU, QIYONG; WEI, ZHENGWEI; ZENG, YANGBIN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 058171/0532 →
Priority Claims (1)
CN 202111122227.X · Sep 24, 2021 · national
Continuity (1)
Related Publication 20230095526A1 · Mar 30, 2023
References Cited (17)
US 8078463B2 · Wasserblat · 2011 [cited by examiner]
US 11217251B2 · York · 2022 [cited by examiner]
US 11423906B2 · Xu · 2022 [cited by examiner]
US 11688412B2 · Zhang · 2023 [cited by examiner]
US 20210217182A1 · Li · 2021 [cited by examiner]
US 20210280171A1 · Phatak · 2021 [cited by examiner]
US 20210326421A1 · Khoury · 2021 [cited by examiner]
US 20220301573A1 · Wang · 2022 [cited by examiner]
US 20220366927A1 · Pishehvar · 2022 [cited by examiner]
US 20220383879A1 · Agarwal · 2022 [cited by examiner]
US 20230068798A1 · Etchart · 2023 [cited by examiner]
CN 112331181A · 2021 [cited by examiner]
CN 112562693A · 2021 [cited by examiner]
Zicheng Liu, Zhengyou Zhang, Li-Wei He, Phil Chou, Energy-Based Sound Source Localization and Gain Normalization for AD HOC Microphone Arrays, ICASSP, 2007, vol. II, pp. 761-764 (Year: 2007). [cited by examiner]
Wei Rao, Chenglin Xu, Eng Siong Chng, Haizhou Li, Target Speaker Extraction for Overlapped Multi-Talker Speaker Verification, arXiv: 1902.02546v1 [eess.AS] Feb. 7, 2019 (Year: 2019). [cited by examiner]
Luo, Yi, and Nima Mesgarani. “Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation.” IEEE/ACM transactions on audio, speech, and language processing 27, No. 8 (2019): 1256-1266. [cited by applicant]
EP International Search Report and Written Opinion for PCT/US2022/044616 malled Jan. 20, 2023. [cited by applicant]
Cited By (2)
US 12,444,429 US 12,670,922