IP Library › Granted Patent US 11,151,996
Granted Patent B2
US 11,151,996 · App. 16/385,630 · Granted Oct 19, 2021

Vocal recognition using generally available speech-to-text systems and user-defined vocal training

Inventors: George A. Saon (Stamford, CT); Nicolò Sgobba (Brno, CZ); Antonello Izzi (Brno, CZ); Erik Rueger (Ockenheim, DE)
Assignee: International Business Machines Corporation
G10L15/22G10L15/063G10L15/07G10L15/26G10L2015/0633G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,151,996
App. No.
16/385,630
Filed
Apr 16, 2019
Granted
Oct 19, 2021
Kind
B2
Examiner
HAN, QI
Art Unit
2659
USPC
704/275
Abstract

Techniques for augmenting the output of generally available speech-to-text systems using local profiles are presented. An example method includes receiving an audio recording of a natural language command. The received audio recording of the natural language command is transmitted to a speech-to-text system, and a text string generated from the audio recording is received from the speech-to-text system. The text string is corrected based on a local profile mapping incorrectly transcribed words from the speech-to-text system to corrected words. A function in a software application is invoked based on the corrected text string.

Claims (61)

1. A method for invoking an action on a computing device using natural language commands, comprising:

generating a local profile by mapping one or more features associated with a user of the computing device to a data set of conversation context for a previous audio recording of a user of the computing device, a voice transcript generated by the speech-to-text system for the previous audio recording, and a corrected textual representation of the previous audio recording;

training a machine learning model based on the mapping of one or more features associated with the user of the computing device to the data set, wherein the mapping of incorrectly parsed words to corrected words comprises a mapping of an identified corrected utterance to incorrectly parsed words in the text string for the received natural language command based on the one or more features associated with the user of the computing device and a conversational context associated with the received natural language command;

receiving an audio recording of a natural language command;

transmitting, to a speech-to-text system, the received audio recording;

receiving, from the speech-to-text system, a text string generated from the received audio recording;

correcting the text string based on the trained machine learning model; and

invoking a function in a software application based on the corrected text string.

2. The method of claim 1 , further comprising:

generating the local profile based on a default local profile associated with one or more features associated with a user of the computing device.

3. The method of claim 2 , wherein the one or more features associated with the user of the computing device comprises features indicative of a probable inflection with which the user of the computing device speaks.

4. The method of claim 1 , further comprising:

outputting the corrected text string for user evaluation; and

receiving a second corrected text string, wherein the function is invoked using the second corrected text string rather than the corrected text string.

5. The method of claim 4 , further comprising:

upon receiving the second corrected text string, adding the text string generated from the received natural language command and the second corrected text string to a data set used to train a machine learning model to identify corrected utterances for the received natural language command based on one or more features associated with a user of a computing device.

6. The method of claim 5 , wherein the machine learning model comprises a classifier trained using unsupervised learning techniques.

7. The method of claim 5 , further comprising:

invoking a training process for a local profile associated with the user of the computing device to update the local profile associated with the user; and

transmitting the updated local profile to the computing device.

8. A system, comprising:

a processor; and

a memory having instructions stored thereon which, when executed by the processor, perform an operation for invoking an action on a computing device using natural language commands, the operation comprising:

generating a local profile by mapping one or more features associated with a user of the computing device to a data set of conversation context for a previous audio recording of a user of the computing device, a voice transcript generated by the speech-to-text system for the previous audio recording, and a corrected textual representation of the previous audio recording;

training a machine learning model based on the mapping of one or more features associated with the user of the computing device to the data set, wherein the mapping of incorrectly parsed words to corrected words comprises a mapping of an identified corrected utterance to incorrectly parsed words in the text string for the received natural language command based on the one or more features associated with the user of the computing device and a conversational context associated with the received natural language command;

receiving an audio recording of a natural language command;

transmitting, to a speech-to-text system, the received audio recording;

receiving, from the speech-to-text system, a text string generated from the received audio recording;

correcting the text string based on the trained machine learning model; and

invoking a function in a software application based on the corrected text string.

9. The system of claim 8 , wherein the operation further comprises:

generating the local profile based on a default local profile associated with one or more features associated with a user of the computing device.

10. The system of claim 9 , wherein the one or more features associated with the user of the computing device comprises features indicative of a probable inflection with which the user of the computing device speaks.

11. The system of claim 8 , wherein the operation further comprises:

outputting the corrected text string for user evaluation; and

receiving a second corrected text string, wherein the function is invoked using the second corrected text string rather than the corrected text string.

12. The system of claim 11 , wherein the operation further comprises:

upon receiving the second corrected text string, adding the text string generated from the received natural language command and the second corrected text string to a data set used to train a machine learning model to identify corrected utterances for the received natural language command based on one or more features associated with a user of a computing device.

13. The system of claim 12 , wherein the machine learning model comprises a classifier trained using unsupervised learning techniques.

14. The system of claim 12 , wherein the operation further comprises:

invoking a training process for a local profile associated with the user of the computing device to update the local profile associated with the user; and

transmitting the updated local profile to the computing device.

15. A non-transitory computer-readable medium having instructions stored thereon which, when executed by one or more processors, performs an operation for invoking an action on a computing device using natural language commands, the operation comprising:

generating a local profile by mapping one or more features associated with a user of the computing device to a data set of conversation context for a previous audio recording of a user of the computing device, a voice transcript generated by the speech-to-text system for the previous audio recording, and a corrected textual representation of the previous audio recording;

training a machine learning model based on the mapping of one or more features associated with the user of the computing device to the data set, wherein the mapping of incorrectly parsed words to corrected words comprises a mapping of an identified corrected utterance to incorrectly parsed words in the text string for the received natural language command based on the one or more features associated with the user of the computing device and a conversational context associated with the received natural language command;

receiving an audio recording of a natural language command;

transmitting, to a speech-to-text system, the received audio recording;

receiving, from the speech-to-text system, a text string generated from the received audio recording;

correcting the text string based on the trained machine learning model; and

invoking a function in a software application based on the corrected text string.

16. The non-transitory computer-readable medium of claim 15 , wherein the operation further comprises:

generating the local profile based on a default local profile associated with one or more features associated with a user of the computing device.

17. The non-transitory computer-readable medium of claim 16 , wherein the one or more features associated with the user of the computing device comprises features indicative of a probable inflection with which the user of the computing device speaks.

18. The non-transitory computer-readable medium of claim 15 , wherein the operation further comprises:

outputting the corrected text string for user evaluation;

receiving a second corrected text string, wherein the function is invoked using the second corrected text string rather than the corrected text string; and

upon receiving the second corrected text string, adding the text string generated from the received natural language command and the second corrected text string to a data set used to train a machine learning model to identify corrected utterances for the received natural language command based on one or more features associated with a user of a computing device.

19. The non-transitory computer-readable medium of claim 18 , wherein the machine learning model comprises a classifier trained using unsupervised learning techniques.

20. The non-transitory computer-readable medium of claim 18 , wherein the operation further comprises:

invoking a training process for a local profile associated with the user of the computing device to update the local profile associated with the user; and

transmitting the updated local profile to the computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2019
From: SAON, GEORGE A; SGOBBA, NICOLO'; IZZI, ANTONELLO; RUEGER, ERIK
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048897/0612 →
Continuity (1)
Related Publication 20200335100A1 · Oct 22, 2020