IP Library Granted Patent US 11,636,846
Granted Patent B2
US 11,636,846 · App. 17/245,019 · Granted Apr 25, 2023

Speech endpointing based on word comparisons

Inventors: Michael Buchanan (Palo Alto, CA); Pravir Kumar Gupta (Los Altos, CA); Christopher Bo Tandiono (Fremont, CA)
Assignee: Google LLC
G10L15/05G10L15/04G10L15/22G10L15/26G10L17/06G10L25/51G10L25/87G10L25/78G10L25/90G10L2015/088G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,636,846
App. No.
17/245,019
Granted
Apr 25, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech endpointing based on word comparisons are described. In one aspect, a method includes the actions of obtaining a transcription of an utterance. The actions further include determining, as a first value, a quantity of text samples in a collection of text samples that (i) include terms that match the transcription, and (ii) do not include any additional terms. The actions further include determining, as a second value, a quantity of text samples in the collection of text samples that (i) include terms that match the transcription, and (ii) include one or more additional terms. The actions further include classifying the utterance as a likely incomplete utterance or not a likely incomplete utterance based at least on comparing the first value and the second value.

Claims (52)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance comprising a sequence of terms;

processing, using a speech recognizer, the audio data to generate a transcription of the utterance by recognizing each term in the sequence of terms;

recording an amount of time after each term is recognized by the speech recognizer;

identifying an endpoint after a respective one of the terms recognized by the speech recognizer when the recorded amount of time after the respective one of the terms is recognized by the speech recognizer satisfies a first threshold before a subsequent term from the sequence of terms is recognized by the speech recognizer;

subsequent to identifying the endpoint after the respective one of the terms recognized by the speech recognizer, determining that the utterance is likely incomplete based on the transcription of the utterance; and

overriding the endpoint identified after the respective one of the terms recognized by the speech recognizer based on determining that the utterance is likely incomplete.

2. The method of claim 1 , wherein determining that the utterance is likely incomplete comprises:

comparing the transcription of the utterance to a first collection of text samples identified as complete utterances;

comparing the transcription of the utterance to a second collection of text samples identified as incomplete utterances; and

determining that the utterance is likely incomplete based on comparing the transcription of the utterance to the first collection of text samples identified as complete utterances and comparing the transcription of the utterance to the second collection of text samples identified as incomplete utterances.

3. The method of claim 2 , wherein comparing the transcription of the utterance to the first collection of text samples identified as complete utterances comprises:

determining a first quantity of text samples in the first collection that match the transcription of the utterance; and

determining a second quantity of text samples in the second collection that match the transcription of the utterance.

4. The method of claim 2 , wherein comparing the transcription of the utterance to the first collection of text samples identified as complete utterances comprises:

determining whether terms in each text sample in the first collection occur in a same order as terms of the transcription of the utterance; and

determining whether terms in each text sample in the second collection occur in a same order as terms of the transcription of the utterance.

5. The method of claim 1 , wherein the operations further comprise:

subsequent to identifying the endpoint after the respective one of the terms recognized by the speech recognizer, determining that the utterance is likely complete based on the transcription of the utterance; and

based on determining that the utterance is likely complete, determining to designate the endpoint identified after the respective one of the terms recognized by the speech recognizer at an end of the audio data of the utterance.

6. The method of claim 5 , wherein the operations further comprise, based on determining to designate the endpoint identified after the respective one of the terms recognized by the speech recognizer at the end of the audio data of the utterance, deactivating a microphone that detected the utterance.

7. The method of claim 1 , wherein overriding the endpoint identified after the respective one of the terms recognized by the speech recognizer comprises maintaining a microphone that detected the utterance in an active state.

8. The method of claim 7 , wherein maintaining the microphone that detected the utterance in the active state permits the speech recognizer to process additional audio data.

9. The method of claim 1 , wherein the speech recognizer resides on a user device.

10. The method of claim 1 , wherein the speech recognizer resides on a server.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance comprising a sequence of terms;

processing, using a speech recognizer, the audio data to generate a transcription of the utterance by recognizing each term in the sequence of terms;

recording an amount of time after each term is recognized by the speech recognizer;

identifying an endpoint after a respective one of the terms recognized by the speech recognizer when the recorded amount of time after the respective one of the terms is recognized by the speech recognizer satisfies a first threshold before a subsequent term of the sequence of terms is recognized by the speech recognizer;

subsequent to identifying the endpoint after the respective one of the terms recognized by the speech recognizer, determining that the utterance is likely incomplete based on the transcription of the utterance; and

overriding the endpoint identified after the respective one of the terms recognized by the speech recognizer based on determining that the utterance is likely incomplete.

12. The system of claim 11 , wherein determining that the utterance is likely incomplete comprises:

comparing the transcription of the utterance to a first collection of text samples identified as complete utterances;

comparing the transcription of the utterance to a second collection of text samples identified as incomplete utterances; and

determining that the utterance is likely incomplete based on comparing the transcription of the utterance to the first collection of text samples identified as complete utterances and comparing the transcription of the utterance to the second collection of text samples identified as incomplete utterances.

13. The system of claim 12 , wherein comparing the transcription of the utterance to the first collection of text samples identified as complete utterances comprises:

determining a first quantity of text samples in the first collection that match the transcription of the utterance; and

determining a second quantity of text samples in the second collection that match the transcription of the utterance.

14. The system of claim 12 , wherein comparing the transcription of the utterance to the first collection of text samples identified as complete utterances comprises:

determining whether terms in each text sample in the first collection occur in a same order as terms of the transcription of the utterance; and

determining whether terms in each text sample in the second collection occur in a same order as terms of the transcription of the utterance.

15. The system of claim 11 , wherein the operations further comprise:

subsequent to identifying the endpoint after the respective one of the terms recognized by the speech recognizer, determining that the utterance is likely complete based on the transcription of the utterance; and

based on determining that the utterance is likely complete, determining to designate the endpoint identified after the respective one of the terms recognized by the speech recognizer at an end of the audio data of the utterance.

16. The system of claim 15 , wherein the operations further comprise, based on determining to designate the endpoint identified after the respective one of the terms recognized by the speech recognizer at the end of the audio data of the utterance, deactivating a microphone that detected the utterance.

17. The system of claim 11 , wherein overriding the endpoint identified after the respective one of the terms recognized by the speech recognizer comprises maintaining a microphone that detected the utterance in an active state.

18. The system of claim 17 , wherein maintaining the microphone that detected the utterance in the active state permits the speech recognizer to process additional audio data.

19. The system of claim 11 , wherein the speech recognizer resides on a user device.

20. The system of claim 11 , wherein the speech recognizer resides on a server.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2021
From: BUCHANAN, MICHAEL; GUPTA, PRAVIR KUMAR; TANDIONO, CHRISTOPHER BO
To: GOOGLE INC.
Reel/Frame 056091/0262 →
CHANGE OF NAME Recorded Apr 30, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 056106/0502 →
Continuity (6)
Continuation 16540795 · Aug 14, 2019
Continuation 16154875 · Oct 9, 2018
Continuation 15156478 · May 17, 2016
Continuation 14681203 · Apr 8, 2015
Provisional Application 61983025 · Apr 23, 2014
Related Publication 20210248995A1 · Aug 12, 2021