IP Library Granted Patent US 12,190,883
Granted Patent B2
US 12,190,883 · App. 18/329,635 · Granted Jan 7, 2025

Speaker recognition adaptation

Inventor: Zeya Chen (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L15/08G10L17/00G10L17/14G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,883
App. No.
18/329,635
Granted
Jan 7, 2025
Kind
B2
Abstract

Techniques for generating, from first speaker recognition data corresponding to at least a first word, second speaker recognition data corresponding to at least a second word are described. During a speaker recognition enrollment process, a device receives audio data corresponding to one or more prompted spoken inputs comprising the at least first word. Using the prompted spoken input(s), the first speaker recognition data (specific to that least first word) is generated. Sometime thereafter, a user may indicate that speaker recognition processing is to be performed using at least a second word. Rather than have the user go through the speaker recognition enrollment process a second time, the device (or a system) may apply a transformation model to the first speaker recognition data to generate second speaker recognition data specific to the at least second word.

Claims (56)

1. A computer-implemented method, comprising:

receiving first data corresponding to at least a first user speaking first content;

receiving second data representing a transformation between how at least a second user is known to speak the first content and how the at least second user is known to speak second content; and

using the first data and the second data to configure a machine learning (ML) model to detect the first user speaking the second content.

2. The computer-implemented method of claim 1 , further comprising:

receiving input audio data; and

processing the input audio data using the ML model to determine the input audio data represents speech of the first user.

3. The computer-implemented method of claim 2 , wherein:

configuration of the ML model is performed by a first device; and

processing the input audio data is performed by a second device.

4. The computer-implemented method of claim 1 , further comprising:

receiving third data corresponding to the first user speaking the second content,

wherein configuration of the ML model further uses the third data.

5. The computer-implemented method of claim 1 , further comprising:

receiving third data corresponding to at least the first user speaking the second content; and

processing the first data and the third data to determine fourth data representing a transformation between first speech characteristics and second speech characteristics, the first speech characteristics corresponding to the first content and the second speech characteristics corresponding to the second content,

wherein the fourth data corresponds to the ML model.

6. The computer-implemented method of claim 1 , wherein the first data further corresponds to a third user speaking the first content.

7. The computer-implemented method of claim 1 , wherein:

the first content corresponds to a first wakeword; and

the second content corresponds to a second wakeword different from the first wakeword.

8. The computer-implemented method of claim 1 , wherein the first data corresponds to at least one quality of frequency domain audio data representing the first user speaking the first content.

9. The computer-implemented method of claim 8 , further comprising:

processing the frequency domain audio data using a feature extraction component to determine the first data.

10. The computer-implemented method of claim 1 , further comprising:

receiving annotation data corresponding to the first data,

wherein configuration of the ML model further uses the annotation data.

11. A system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive first data corresponding to at least a first user speaking first content;

receive second data representing a transformation between how at least a second user is known to speak the first content and how the at least second user is known to speak second content; and

use the first data and the second data to configure a machine learning (ML) model to detect the first user speaking the second content.

12. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive input audio data; and

process the input audio data using the ML model to determine the input audio data represents speech of the first user.

13. The system of claim 12 , wherein:

configuration of the ML model is performed by a first device; and

processing the input audio data is performed by a second device.

14. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive third data corresponding to the first user speaking the second content,

wherein configuration of the ML model further uses the third data.

15. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive third data corresponding to at least the first user speaking the second content; and

process the first data and the third data to determine fourth data representing a transformation between first speech characteristics and second speech characteristics, the first speech characteristics corresponding to the first content and the second speech characteristics corresponding to the second content,

wherein the fourth data corresponds to the ML model.

16. The system of claim 11 , wherein the first data further corresponds to a third user speaking the first content.

17. The system of claim 11 , wherein:

the first content corresponds to a first wakeword; and

the second content corresponds to a second wakeword different from the first wakeword.

18. The system of claim 11 , wherein the first data corresponds to at least one quality of frequency domain audio data representing the first user speaking the first content.

19. The system of claim 18 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

process the frequency domain audio data using a feature extraction component to determine the first data.

20. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive annotation data corresponding to the first data,

wherein configuration of the ML model further uses the annotation data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2023
From: CHEN, ZEYA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 063862/0908 →
Continuity (2)
Continuation 16912119 · Jun 25, 2020
Related Publication 20240013784A1 · Jan 11, 2024
References Cited (5)
US 11355102B1 · Mishchenko · 2022 [cited by examiner]
US 20170372694A1 · Ushio · 2017 [cited by examiner]
US 20190156835A1 · Church · 2019 [cited by examiner]
US 20190341057A1 · Zhang · 2019 [cited by examiner]
US 20210183367A1 · Sharifi · 2021 [cited by examiner]