IP Library › Granted Patent US 12,347,439
Granted Patent B2
US 12,347,439 · App. 18/153,932 · Granted Jul 1, 2025

Multi-task learning for personalized keyword spotting

Inventors: Seunghan Yang (Incheon, KR); Byeonggeun Kim (Seoul, KR); Inseop Chung (Seoul, KR); Simyung Chang (Suwon, KR)
Assignee: QUALCOMM Incorporated
G10L17/14G10L17/04G10L17/18G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,439
App. No.
18/153,932
Granted
Jul 1, 2025
Kind
B2
Abstract

Systems and techniques are provided for processing audio data. For example, the systems and techniques can be used for personalized keyword spotting through multi-task learning (PK-MTL). A process can include obtaining an audio sample, generating a representation of a keyword based on the audio sample, and generating a representation of a speaker based on the audio sample. The speaker can be associated with the keyword. A first similarity score can be determined based on a reference representation and one or more of the representation of the keyword and a representation of the speaker. The reference representation can be associated with one or more of the keyword and the speaker. A keyword spotting (KWS) output can be generated based on analyzing the first similarity score against at least a first threshold, wherein the KWS output accepts or rejects the audio sample as including a target keyword.

Claims (62)

1. A method for processing one or more audio samples, comprising:

obtaining an audio sample;

generating a representation of a keyword, wherein the representation of the keyword is generated based on the audio sample;

generating a representation of a speaker, wherein the speaker is associated with the keyword and the representation of the speaker is generated based on the audio sample;

determining a first similarity score based on a reference representation and one or more of the representation of the keyword and the representation of the speaker, wherein the reference representation is associated with one or more of the keyword and the speaker; and

generating a keyword spotting (KWS) output based on analyzing the first similarity score against at least a first threshold, wherein the KWS output accepts or rejects the audio sample as including a target keyword;

wherein determining the first similarity score comprises:

generating a keyword similarity score between the representation of the keyword and a reference representation of the target keyword;

generating a speaker similarity score between the representation of the speaker and a reference representation of the speaker; and

determining the first similarity score as a combined similarity score generated based at least in part on the keyword similarity score and the speaker similarity score.

2. The method of claim 1 , wherein:

the representation of the keyword is a keyword embedding generated based on the audio sample; and

the representation of the speaker is a speaker embedding generated based on the audio sample.

3. The method of claim 1 , wherein one or more of the representation of the keyword or the representation of the speaker is generated using a multi-task learning (MTL) machine learning network.

4. The method of claim 1 , wherein the first similarity score and the KWS output are generated using a task-adaptation machine learning network.

5. The method of claim 1 , wherein one or more of the keyword similarity score or the speaker similarity score is a cosine similarity score.

6. The method of claim 1 , wherein the combined similarity score is generated using a score combination function.

7. The method of claim 6 , wherein:

the score combination function is a linear combination function between the keyword similarity score and the speaker similarity score; and

the linear combination function includes at least a first tunable weighting parameter.

8. The method of claim 7 , wherein the score combination function is associated with one or more neural networks and is trained to tune the first tunable weighting parameter to minimize a keyword spotting (KWS) false rejection rate (FRR) associated with the one or more neural networks.

9. The method of claim 7 , further comprising:

setting the first tunable weighting parameter to a first value to perform target user-biased keyword spotting (TB-KWS); and

setting the first tunable weighting parameter to a second value to perform target user-only keyword spotting (TO-KWS), wherein the first value is larger than the second value.

10. The method of claim 1 , wherein determining the first similarity score comprises:

generating a task-specific embedding using the representation of the keyword and the representation of the speaker; and

determining the first similarity score as a cosine similarity score between the task-specific embedding and the reference representation, wherein the reference representation is a learnable weight for keyword classification.

11. The method of claim 10 , further comprising generating the task-specific embedding based on an output of a first neural network, the output of the first neural network including a target user-biased keyword spotting (TB-KWS) task-specific embedding.

12. The method of claim 11 , further comprising generating the task-specific embedding based on an output of a second neural network, the output of the second neural network including a target user-only keyword spotting (TO-KWS) task-specific embedding.

13. An apparatus for processing one or more audio samples, comprising:

at least one memory; and

at least one processor coupled to the at least one memory, the at least one processor configured to:

obtain an audio sample;

generate a representation of a keyword, wherein the representation of the keyword is generated based on the audio sample;

generate a representation of a speaker, wherein the speaker is associated with the keyword and the representation of the speaker is generated based on the audio sample;

determine a first similarity score based on a reference representation and one or more of the representation of the keyword and the representation of the speaker, wherein the reference representation is associated with one or more of the keyword and the speaker; and

generate a keyword spotting (KWS) output based on analyzing the first similarity score against at least a first threshold, wherein the KWS output accepts or rejects the audio sample as including a target keyword;

wherein, to determine the first similarity score, the at least one processor is configured to:

generate a keyword similarity score between the representation of the keyword and a reference representation of the target keyword;

generate a speaker similarity score between the representation of the speaker and a reference representation of the speaker; and

determine the first similarity score as a combined similarity score generated based at least in part on the keyword similarity score and the speaker similarity score.

14. The apparatus of claim 13 , wherein:

the representation of the keyword is a keyword embedding generated based on the audio sample; and

the representation of the speaker is a speaker embedding generated based on the audio sample.

15. The apparatus of claim 13 , wherein one or more of the representation of the keyword or the representation of the speaker is generated using a multi-task learning (MTL) machine learning network.

16. The apparatus of claim 13 , wherein the first similarity score and the KWS output are generated using a task-adaptation machine learning network.

17. The apparatus of claim 13 , wherein one or more of the keyword similarity score or the speaker similarity score is a cosine similarity score.

18. The apparatus of claim 13 , wherein the combined similarity score is generated using a score combination function.

19. The apparatus of claim 18 , wherein:

the score combination function is a linear combination function between the keyword similarity score and the speaker similarity score; and

the linear combination function includes at least a first tunable weighting parameter.

20. The apparatus of claim 19 , wherein the score combination function is associated with one or more neural networks and is trained to tune the first tunable weighting parameter to minimize a keyword spotting (KWS) false rejection rate (FRR) associated with the one or more neural networks.

21. The apparatus of claim 19 , wherein the at least one processor is further configured to:

set the first tunable weighting parameter to a first value to perform target user-biased keyword spotting (TB-KWS); and

set the first tunable weighting parameter to a second value to perform target user-only keyword spotting (TO-KWS), wherein the first value is larger than the second value.

22. The apparatus of claim 13 , wherein, to determine the first similarity score, the at least one processor is configured to:

generate a task-specific embedding using the representation of the keyword and the representation of the speaker; and

determine the first similarity score as a cosine similarity score between the task-specific embedding and the reference representation, wherein the reference representation is a learnable weight for keyword classification.

23. The apparatus of claim 22 , wherein the at least one processor is further configured to:

generate the task-specific embedding based on an output of a first neural network, the output of the first neural network including a target user-biased keyword spotting (TB-KWS) task-specific embedding.

24. The apparatus of claim 23 , wherein the at least one processor is further configured to:

generate the task-specific embedding based on an output of a second neural network, the output of the second neural network including a target user-only keyword spotting (TO-KWS) task-specific embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2023
From: YANG, SEUNGHAN; KIM, BYEONGGEUN; CHUNG, INSEOP; CHANG, SIMYUNG
To: QUALCOMM INCORPORATED
Reel/Frame 062861/0639 →
Continuity (2)
Provisional Application 63322164 · Mar 21, 2022
Related Publication 20230298592A1 · Sep 21, 2023
References Cited (9)
US 9711148B1 · Sharifi · 2017 [cited by examiner]
US 20140122087A1 · Macho · 2014 [cited by examiner]
US 20180286433A1 · Hicks · 2018 [cited by examiner]
US 20190043507A1 · Huang et al. · 2019 [cited by applicant]
US 20190392839A1 · Fujimura · 2019 [cited by examiner]
US 20220157329A1 · Choi · 2022 [cited by examiner]
US 20220261218A1 · Shin · 2022 [cited by examiner]
International Search Report and Written Opinion—PCT/US2023/060959—ISA/EPO—Apr. 11, 2023. [cited by applicant]
Rikhye R., et al., “Multi-User Voicefilter-Lite via Attentive Speaker Embedding”, 2021 IEEE Automatic Speech Recognition and Understanding Workshop , IEEE, Dec. 13, 2021, XP034076961, pp. 275-282, Abstract, paragraphs 1… [cited by applicant]