IP Library Patent Application 19010299
Patent Application
App. No. 19/010,299

Disfluency Detection Models for Natural Conversational Voice Systems

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/010,299
Abstract

A method includes receiving a sequence of acoustic frames characterizing one or more utterances. At each of a plurality of output steps, the method also includes generating, by an encoder network of a speech recognition model, a higher order feature representation for a corresponding acoustic frame of the sequence of acoustic frames, generating, by a prediction network of the speech recognition model, a hidden representation for a corresponding sequence of non-blank symbols output by a final softmax layer of the speech recognition model, and generating, by a first joint network of the speech recognition model that receives the higher order feature representation generated by the encoder network and the dense representation generated by the prediction network, a probability distribution that the corresponding time step corresponds to a pause and an end of speech.

Claims (50)

1 . A computer-implemented method executing on data processing hardware that causes the data processing hardware to perform operations comprising:

as a user speaks an utterance directed toward a digital assistant application and captured by a microphone of a user device:

receiving a first sequence of acoustic frames characterizing a first portion of the utterance spoken by the user;

processing, using a speech recognition model, the first sequence of acoustic frames to generate first speech recognition results for the first portion of the utterance;

after receiving the first sequence of acoustic frames, receiving a second sequence of acoustic frames characterizing a second portion of the utterance spoken by the user;

processing, using the speech recognition model, the second sequence of acoustic frames to detect a presence of a disfluency in the second portion of the utterance;

based on detecting the presence of the disfluency in the second portion of the utterance, determining that the user has not finished speaking the utterance;

after receiving the second sequence of acoustic frames, receiving a third sequence of acoustic frames characterizing a third portion of the utterance spoken by the user; and

processing, using the speech recognition model, the third sequence of acoustic frames to:

generate second speech recognition results for the third portion of the utterance; and

detect an end of speech event at an end of the third portion of the utterance; and

based on detecting the end of speech event, generating a response to the utterance.

2 . The computer-implemented method of claim 1 , wherein detecting the presence of the disfluency in the second portion of the utterance comprises detecting a presence of a filler word spoken by the user in the second portion of the utterance.

3 . The computer-implemented method of claim 2 , wherein processing the second sequence of acoustic frames to detect the presence of the disfluency further comprises processing, using the speech recognition model, the second sequence of acoustic frames to generate third speech recognition results for the second portion of the utterance, the third speech recognition results comprising a transcript of the filler word spoken by the user.

4 . The computer-implemented method of claim 3 , wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device, a transcription of the utterance that comprises the first speech recognition results generated for the first portion of the utterance, the third speech recognition results generated for the second portion of the utterance, and the second speech recognition results generated for the third portion of the utterance.

5 . The computer-implemented method of claim 1 , wherein detecting the presence of the disfluency in the second portion of the utterance comprises detecting a presence of a pause in speech in the second portion of the utterance.

6 . The computer-implemented method of claim 1 , wherein the operations further comprise, based on detecting the end of speech event, processing the first speech recognition results and the second speech recognition results to execute a query specified by the utterance spoken by the user.

7 . The computer-implemented method of claim 6 , wherein the operations further comprise providing, for audible output from the user device, a synthesized speech representation of the response to the query.

8 . The computer-implemented method of claim 6 , wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device:

the first speech recognition results for the first portion of the utterance;

the second speech recognition results for the second portion of the utterance; and

the response to the query.

9 . The computer-implemented method of claim 1 , wherein the operations further comprise, based on detecting the presence of the disfluency in the second portion of the utterance, generating an acknowledgement response, the acknowledgement response indicating to the user that the digital assistant application is waiting for the user to finish speaking the utterance.

10 . The computer-implemented method of claim 1 , wherein the speech recognition model comprises a stack of self-attention blocks.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

as a user speaks an utterance directed toward a digital assistant application and captured by a microphone of a user device:

receiving a first sequence of acoustic frames characterizing a first portion of the utterance spoken by the user;

processing, using a speech recognition model, the first sequence of acoustic frames to generate first speech recognition results for the first portion of the utterance;

after receiving the first sequence of acoustic frames, receiving a second sequence of acoustic frames characterizing a second portion of the utterance spoken by the user;

processing, using the speech recognition model, the second sequence of acoustic frames to detect a presence of a disfluency in the second portion of the utterance;

based on detecting the presence of the disfluency in the second portion of the utterance, determining that the user has not finished speaking the utterance;

after receiving the second sequence of acoustic frames, receiving a third sequence of acoustic frames characterizing a third portion of the utterance spoken by the user; and

processing, using the speech recognition model, the third sequence of acoustic frames to:

generate second speech recognition results for the third portion of the utterance; and

detect an end of speech event at an end of the third portion of the utterance; and

based on detecting the end of speech event, triggering a microphone closing event by the user device.

12 . The system of claim 11 , wherein detecting the presence of the disfluency in the second portion of the utterance comprises detecting a presence of a filler word spoken by the user in the second portion of the utterance.

13 . The system of claim 12 , wherein processing the second sequence of acoustic frames to detect the presence of the disfluency further comprises processing, using the speech recognition model, the second sequence of acoustic frames to generate third speech recognition results for the second portion of the utterance, the third speech recognition results comprising a transcript of the filler word spoken by the user.

14 . The system of claim 13 , wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device, a transcription of the utterance that comprises the first speech recognition results generated for the first portion of the utterance, the third speech recognition results generated for the second portion of the utterance, and the second speech recognition results generated for the third portion of the utterance.

15 . The system of claim 11 , wherein detecting the presence of the disfluency in the second portion of the utterance comprises detecting a presence of a pause in speech in the second portion of the utterance.

16 . The system of claim 11 , wherein the operations further comprise, based on detecting the end of speech event, processing the first speech recognition results and the second speech recognition results to execute a query specified by the utterance spoken by the user.

17 . The system of claim 16 , wherein the operations further comprise providing, for audible output from the user device, a synthesized speech representation of the response to the query.

18 . The system of claim 16 , wherein the operations further comprise providing, for display in a graphical user interface displayed on a screen of the user device:

the first speech recognition results for the first portion of the utterance;

the second speech recognition results for the second portion of the utterance; and

the response to the query.

19 . The system of claim 11 , wherein the operations further comprise, based on detecting the presence of the disfluency in the second portion of the utterance, generating an acknowledgement response, the acknowledgement response indicating to the user that the digital assistant application is waiting for the user to finish speaking the utterance.

20 . The system of claim 11 , wherein the speech recognition model comprises a stack of self-attention blocks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2025
From: CHANG, SHUO-YIIN; LI, BO; SAINATH, TARA N.; STROHMAN, TREVOR; ZHANG, CHAO
To: GOOGLE LLC
Reel/Frame 069749/0350 →