IP Library Granted Patent US 10,504,505
Granted Patent B2
US 10,504,505 · App. 15/830,511 · Granted Dec 10, 2019

System and method for speech personalization by need

Inventors: Andrej Ljolje (Morris Plains, NJ); Alistair D. Conkie (San Jose, CA); Ann K. Syrdal (San Jose, CA)
Assignee: NUANCE COMMUNICATIONS, INC.
G10L15/07G10L15/10G10L15/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,504,505
App. No.
15/830,511
Granted
Dec 10, 2019
Kind
B2
Abstract

Disclosed herein are systems, computer-implemented methods, and tangible computer-readable storage media for speaker recognition personalization. The method recognizes speech received from a speaker interacting with a speech interface using a set of allocated resources, the set of allocated resources including bandwidth, processor time, memory, and storage. The method records metrics associated with the recognized speech, and after recording the metrics, modifies at least one of the allocated resources in the set of allocated resources commensurate with the recorded metrics. The method recognizes additional speech from the speaker using the modified set of allocated resources. Metrics can include a speech recognition confidence score, processing speed, dialog behavior, requests for repeats, negative responses to confirmations, and task completions. The method can further store a speaker personalization profile having information for the modified set of allocated resources and recognize speech associated with the speaker based on the speaker personalization profile.

Claims (31)

1. A method comprising:

recognizing speech for each of a plurality of speakers;

while processing further speech from the each speaker of the plurality of speakers, modifying, via a processor, an allocation of computer resources of a speech interface based on a metric gathered from the recognizing of the speech for each of the plurality of speakers, the metric associated with at least one of a request for repetition, a negative response to confirmation, and a task completion, to yield a modified speech interface; and

recognizing, via the modified speech interface, additional speech.

2. The method of claim 1 , wherein the speech interface utilizes at least one of bandwidth and processor time to allocate resources for the recognizing speech.

3. The method of claim 1 , wherein the metric further comprises a speech recognition confidence score, a processing speed, and a dialog behavior.

4. The method of claim 1 , wherein an identified speaker was determined to be frustrated and have great difficulty in a prior session.

5. The method of claim 1 , further comprising storing a speaker personalization profile having information for the modified speech interface.

6. The method of claim 5 , further comprising recognizing speech associated with an identified speaker based on the speaker personalization profile.

7. The method of claim 5 , further comprising storing the speaker personalization profile on a personalization server storing multiple speaker personalization profiles.

8. The method of claim 5 , wherein multiple speakers are associated with the speaker personalization profile.

9. The method of claim 1 , wherein the modified speech interface is associated with a class of similar speakers.

10. The method of claim 1 , wherein modifying the allocation of resources is based on a difficulty threshold associated with how well the speaker interacts with the speech interface.

11. The method of claim 1 , further comprising progressively applying the modified speech interface.

12. The method of claim 1 , wherein the resources each comprise at least one of memory and storage.

13. The method of claim 1 , wherein an allocation of resources in the modified speech interface is greater than its corresponding allocation in a set of allocated resources prior to the modifying.

14. The method of claim 1 , wherein an allocation of resources in the modified speech interface is less than its corresponding allocation in a set of allocated resources prior to the modifying.

15. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, result in the processor performing operations comprising:

recognizing speech for each of a plurality of speakers;

while processing further speech from the each speaker of the plurality of speakers, modifying an allocation of computer resources of a speech interface based on metric gathered from the recognizing of the speech for each of the plurality of speakers, the metric associated with at least one of a request for repetition, a negative response to confirmation, and a task completion, to yield a modified speech interface; and

recognizing, via the modified speech interface, additional speech.

16. The system of claim 15 , wherein the speech interface utilizes at least one of bandwidth and processor time to allocate resources for the recognizing speech.

17. The system of claim 15 , wherein the metric further comprises a speech recognition confidence score, a processing speed, and a dialog behavior.

18. The system of claim 15 , wherein the identified speaker was determined to be frustrated and have great difficulty in a prior session.

19. The system of claim 15 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising storing a speaker personalization profile having information for the modified speech interface.

20. A computer-readable storage device having instructions stored which, when executed by a computing device, result in the computing device performing operations comprising:

recognizing speech for each of a plurality of speakers;

while processing further speech from the each speaker of the plurality of speakers, modifying an allocation of computer resources of a speech interface based on metric gathered from the recognizing of the speech for each of the plurality of speakers, the metric associated with at least one of a request for repetition, a negative response to confirmation, and a task completion, to yield a modified speech interface; and

recognizing, via the modified speech interface, additional speech.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065225/0295 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2019
From: LJOLJE, ANDREJ; CONKIE, ALISTAIR D.; SYRDAL, ANN K.
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 050780/0336 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2019
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 050780/0620 →
Continuity (3)
Continuation 14679508 · Apr 6, 2015
Continuation 12480864 · Jun 9, 2009
Related Publication 20180090129A1 · Mar 29, 2018