IP Library Granted Patent US 9,875,741
Granted Patent B2
US 9,875,741 · App. 14/775,729 · Granted Jan 23, 2018

Selective speech recognition for chat and digital personal assistant systems

Inventors: Ilya Genadevich Gelfenbeyn (Sunnyvale, CA); Artem Goncharuk (Arlington, VA); Ilya Andreevich Platonov (Berdsk, RU); Pavel Aleksandrovich Sirotin (Sunnyvale, CA); Olga Aleksandrovna Gelfenbeyn (Yurga, RU)
Assignee: GOOGLE LLC
G10L15/32G10L15/02G10L15/07G10L15/22G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,875,741
App. No.
14/775,729
Granted
Jan 23, 2018
Kind
B2
Abstract

Disclosed are computer-implemented methods and systems for dynamic selection of speech recognition systems for the use in Chat Information Systems (CIS) based on multiple criteria and context of human-machine interaction. Specifically, once a first user audio input is received, it is analyzed so as to locate specific triggers, determine the context of the interaction or predict the subsequent user audio inputs. Based on at least one of these criteria, one of a free-diction recognizer, pattern-based recognizer, address book based recognizer or dynamically created recognizer is selected for recognizing the subsequent user audio input. The methods described herein increase the accuracy of automatic recognition of user voice commands, thereby enhancing overall user experience of using CIS, chat agents and similar digital personal assistant systems.

Claims (25)

1. A method for speech recognition in a chat information system (CIS), the method comprising:

receiving, by a processor operatively coupled to a memory, an audio input;

separating, by the processor, the audio input into a plurality of parts having at least a first part of the audio input and a second part of the audio input;

selecting, from a plurality of speech recognizers, a specific first speech recognizer to recognize the first part of the audio input, wherein selecting of the specific first speech recognizer to recognize the first part of the audio input is by the processor and is based on predetermined criteria,

wherein each of the plurality of speech recognizers, from which the specific first speech recognizer is selected to recognize the first part of the audio input, is configured to generate, based on a corresponding audio input, a plurality of outputs provided with corresponding confidence levels;

recognizing, by the specific first speech recognizer of a plurality of speech recognizers, the first part of the audio input to generate a first recognized input;

analyzing, by the processor, the first recognized input associated with the first part of the audio input to identify at least one first trigger in the first recognized input;

predicting, by the processor, a type of the second part of the audio input based at least in part on the at least one first trigger;

based on the prediction of the type of the second part of the audio input, selecting, by the processor, a specific second speech recognizer from the plurality of speech recognizers;

recognizing, by the specific second speech recognizer, the second part of the audio input to generate a second recognized input;

analyzing, by the processor, the second recognized input to identify at least one second trigger in the second recognized input;

predicting, by the processor, types of further parts of the audio input based at least in part on triggers identified in recognized inputs;

selecting, from the plurality of speech recognizers, further specific speech recognizers based on the predicted types of the further parts of the audio input, the further specific speech recognizers being in addition to the first speech recognizer and the second speech recognizer; and

recognizing, by the further specific speech recognizers, the further parts of the audio input until all parts of the audio input are recognized.

2. The method of claim 1 , wherein the separating of the audio input comprises recognizing, by one of the plurality of speech recognizers, at least a beginning part of the audio input to generate a recognized input.

3. The method of claim 2 , further comprising selecting, by the processor, the specific first speech recognizer based at least in part on the recognized input.

4. The method of claim 1 , wherein the at least one first trigger includes a type of the audio input identified based at least in part on the first recognized input.

5. The method of claim 4 , wherein the type of the audio input includes a free speech input or a pattern-based speech input.

6. The method of claim 5 , wherein the pattern-based speech input includes at least one of the following: a name, a nickname, a title, an address, and a number.

7. The method of claim 1 , wherein the specific first speech recognizer or the specific second speech recognizer includes a pattern-based speech recognizer.

8. The method of claim 1 , wherein the specific first speech recognizer or the specific second speech recognizer includes a free-dictation recognizer.

9. The method of claim 1 , wherein the specific first speech recognizer or the specific second speech recognizer includes an address book based recognizer.

10. The method of claim 1 , wherein the specific first speech recognizer or the specific second speech recognizer includes a dynamically created recognizer.

11. The method of claim 1 , further comprising combining, by the processor, the first recognized input and the second recognized input.

12. The method of claim 1 , further comprising generating, by the CIS, a response based at least in part on the first recognized input or the second recognized input.

Assignments (3)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2017
From: OOO SPEAKTOIT LLC
To: GOOGLE INC.
Reel/Frame 041263/0894 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2015
From: GELFENBEYN, ILYA GENADEVICH; GONCHARUK, ARTEM; PLATONOV, ILYA ANDREEVICH; SIROTIN, PAVEL ALEKSANDROVICH; GELFENBEYN, OLGA ALEKSANDROVNA
To: OOO "SPEAKTOIT"
Reel/Frame 036551/0110 →
Continuity (1)
Related Publication 20160027440A1 · Jan 28, 2016