IP Library Granted Patent US 9,607,613
Granted Patent B2
US 9,607,613 · App. 14/681,203 · Granted Mar 28, 2017

Speech endpointing based on word comparisons

Inventors: Michael Buchanan (Mountain View, CA); Pravir Kumar Gupta (Mountain View, CA); Christopher Bo Tandiono (Fremont, CA)
Assignee: Google Inc.
G10L15/05G10L15/04G10L15/22G10L15/26G10L25/51G10L25/87G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,607,613
App. No.
14/681,203
Filed
Apr 8, 2015
Granted
Mar 28, 2017
Kind
B2
Art Unit
2675
USPC
704/235
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech endpointing based on word comparisons are described. In one aspect, a method includes the actions of obtaining a transcription of an utterance. The actions further include determining, as a first value, a quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms. The actions further include determining, as a second value, a quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms. The actions further include classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first value and the second value.

Claims (78)

1. A computer-implemented method comprising:

obtaining a transcription of an utterance;

determining a first quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms;

determining a second quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms;

comparing the first quantity and the second quantity;

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance.

2. The method of claim 1 , wherein determining a first quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms comprises:

determining that, in each text sample, that terms that match the transcription occur in a same order as in the transcription.

3. The method of claim 1 , wherein determining a second quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms comprises:

determining that, in each text sample, the terms that match the transcription occur at a prefix of each text sample.

4. The method of claim 1 , wherein:

comparing the first quantity and the second quantity comprises:

determining a ratio of the first quantity to the second quantity; and

determining that the ratio satisfies a threshold ratio; and

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises:

based on determining that the ratio satisfies the threshold ratio, classifying the utterance as a likely incomplete utterance.

5. The method of claim 1 , wherein:

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as a likely incomplete utterance; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance comprises based on classifying the utterance as a likely incomplete utterance, maintaining a microphone in an active state to receive an additional utterance.

6. The method of claim 1 , wherein:

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as not a likely incomplete utterance; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance comprises based on classifying the utterance as not a likely incomplete utterance, deactivating a microphone.

7. The method of claim 1 , comprising:

receiving data indicating that the utterance is complete;

wherein classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as a likely incomplete utterance; and

based on classifying the utterance as a likely incomplete utterance, overriding the data indicating that the utterance is complete.

8. The method of claim 1 , wherein the transcription of the utterance is a search query and the collection of text samples is a collection of search queries.

9. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining a transcription of an utterance;

determining a first quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms;

determining second quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms;

comparing the first quantity and the second quantity;

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating a microphone based on classifying the utterance as not a likely incomplete utterance.

10. The system of claim 9 , wherein determining a first quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms comprises:

determining that, in each text sample, that terms that match the transcription occur in a same order as in the transcription.

11. The system of claim 9 , wherein determining a second quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms comprises:

determining that, in each text sample, the terms that match the transcription occur at a prefix of each text sample.

12. The system of claim 9 , wherein:

comparing the first quantity and the second quantity comprises:

determining a ratio of the first quantity to the second quantity; and

determining that the ratio satisfies a threshold ratio; and

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises:

based on determining that the ratio satisfies the threshold ratio, classifying the utterance as a likely incomplete utterance.

13. The system of claim 9 , wherein:

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as a likely incomplete utterance; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance comprises based on classifying the utterance as a likely incomplete utterance, maintaining a microphone in an active state to receive an additional utterance.

14. The system of claim 9 , wherein:

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as not a likely incomplete utterance; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance comprises based on classifying the utterance as not a likely incomplete utterance, deactivating a microphone.

15. The system of claim 9 , the operations further comprising:

receiving data indicating that the utterance is complete;

wherein classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as a likely incomplete utterance; and

based on classifying the utterance as a likely incomplete utterance, overriding the data indicating that the utterance is complete.

16. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

obtaining a transcription of an utterance;

determining a first quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms;

determining a second quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms;

comparing the first quantity and the second quantity;

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance.

17. The medium of claim 16 , wherein determining a first quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms comprises:

determining that, in each text sample, the terms that match the transcription occur at a prefix of each text sample.

18. The medium of claim 16 , wherein:

comparing the first quantity and the second quantity comprises:

determining a ratio of the first quantity to the second quantity; and

determining that the ratio satisfies a threshold ratio; and

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises:

based on determining that the ratio satisfies the threshold ratio, classifying the utterance as a likely incomplete utterance.

19. The medium of claim 16 , wherein:

classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as a likely incomplete utterance; and

maintaining a microphone in an active state to receive an additional utterance based on classifying the utterance as a likely incomplete, or deactivating the microphone based on classifying the utterance as not a likely incomplete utterance comprises based on classifying the utterance as a likely incomplete utterance, maintaining a microphone in an active state to receive an additional utterance.

20. The medium of claim 16 , the operations further comprising:

receiving data indicating that the utterance is complete;

wherein classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first quantity and the second quantity comprises classifying the utterance as a likely incomplete utterance; and

based on classifying the utterance as a likely incomplete utterance, overriding the data indicating that the utterance is complete.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044097/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2015
From: BUCHANAN, MICHAEL; GUPTA, PRAVIR KUMAR; TANDIONO, CHRISTOPHER BO
To: GOOGLE INC.
Reel/Frame 036355/0128 →
Continuity (2)
Provisional Application 61983025 · Apr 23, 2014
Related Publication 20150310879A1 · Oct 29, 2015