IP Library › Granted Patent US 12,106,758
Granted Patent B2
US 12,106,758 · App. 17/322,765 · Granted Oct 1, 2024

Voice commands for an automated assistant utilized in smart dictation

Inventors: Victor Carbune (Zurich, CH); Alvin Abdagic (Zurich, CH); Behshad Behzadi (Freienbach, CH); Jacopo Sannazzaro Natta (Berkeley, CA); Julia Proskurnia (Zurich, CH); Krzysztof Andrzej Goj (Zurich, CH); Srikanth Pandiri (Zurich, CH); Viesturs Zarins (Zurich, CH); Nicolo D'Ercole (Oberrieden, CH); Zaheed Sabur (Baar, CH); Luv Kothari (Sunnyvale, CA)
Assignee: GOOGLE LLC
G10L15/26G06F3/0488G06N20/00G10L15/18G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,106,758
App. No.
17/322,765
Granted
Oct 1, 2024
Kind
B2
Abstract

Systems and methods described herein relate to determining whether to incorporate recognized text, that corresponds to a spoken utterance of a user of a client device, into a transcription displayed at the client device, or to cause an assistant command, that is associated with the transcription and that is based on the recognized text, to be performed by an automated assistant implemented by the client device. The spoken utterance is received during a dictation session between the user and the automated assistant. Implementations can process, using automatic speech recognition model(s), audio data that captures the spoken utterance to generate the recognized text. Further, implementations can determine whether to incorporate the recognized text into the transcription or cause the assistant command to be performed based on touch input being directed to the transcription, a state of the transcription, and/or audio-based characteristic(s) of the spoken utterance.

Claims (66)

1. A method implemented by one or more processors, the method comprising:

receiving audio data that captures a spoken utterance of a user of a client device, the audio data being generated by one or more microphones of the client device;

determining whether touch input of the user is being simultaneously directed to a transcription, that is displayed at the client device via a software application accessible at the client device, at the same time the audio data that captures the spoken utterance is received;

in response to determining that no touch input of the user is being simultaneously directed to the transcription at the same time the audio data that captures the spoken utterance is received:

determining to incorporate recognized text, that corresponds to the spoken utterance, into the transcription;

in response to determining that touch input of the user is being simultaneously directed to the transcription at the same time the audio data that captures the spoken utterance is received:

determining, based on one or more terms of the spoken utterance, whether to:

incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or

perform an assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance;

in response to determining to incorporate the recognized text that corresponds to the spoken utterance into the transcription:

automatically incorporating the recognized text that corresponds to the spoken utterance into the transcription; and

in response to determining to perform the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance:

causing an automated assistant to perform the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance.

2. The method of claim 1 , further comprising:

processing, using an automatic speech recognition (ASR) model, the audio data that captures the spoken utterance to generate the recognized text that corresponds to the spoken utterance.

3. The method of claim 2 , further comprising:

processing, using a natural language understanding (NLU) model, the recognized text that corresponds to the spoken utterance to generate annotated recognized text.

4. The method of claim 3 , further comprising:

determining the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, wherein determining the assistant command, that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, is based on the annotated recognized text.

5. The method of claim 1 , wherein the touch input of the user is being directed to one or more textual segments of the transcription that is displayed at the client device.

6. The method of claim 5 , wherein the touch input of the user graphically demarcates one or more of the textual segments of the transcription that is displayed at the client device.

7. The method of claim 6 , wherein determining whether to incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or to perform the assistant command, that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, comprises determining to perform the assistant command based on the touch input of the user graphically demarcating one or more of the textual segments of the transcription.

8. The method of claim 1 , wherein the touch input of the user is being directed to one or more fields of the transcription that is displayed at the client device.

9. The method of claim 8 , wherein determining whether to incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or to perform the assistant command, that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, comprises determining to perform the assistant command based on the touch input of the user being directed to one or more fields of the transcription.

10. The method of claim 1 , wherein automatically incorporating the recognized text that corresponds to the spoken utterance into the transcription comprises:

causing the recognized text to be visually displayed to the user via the software application accessible at the client device as part of the transcription.

11. The method of claim 10 , wherein causing the recognized text to be visually displayed to the user via the software application accessible at the client device as part of the transcription comprises causing the recognized text to be maintained in the transcription after additional text is incorporated into the transcription.

12. A system comprising:

at least one processor; and

memory storing instructions that, when executed, cause the at least one processor to be operable to:

receive audio data that captures a spoken utterance of a user of a client device, the audio data being generated by one or more microphones of the client device;

determine whether touch input of the user is being simultaneously directed to a transcription, that is displayed at the client device via a software application accessible at the client device, the audio data that captures the spoken utterance is received;

in response to determining that no touch input of the user is being simultaneously directed to the transcription at the same time the audio data that captures the spoken utterance is received:

determine to incorporate recognized text, that corresponds to the spoken utterance, into the transcription;

in response to determining that touch input of the user is being simultaneously directed to the transcription at the same time the audio data that captures the spoken utterance is received:

determine, based on one or more terms of the spoken utterance, whether to:

incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or

perform an assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance;

in response to determining to incorporate the recognized text that corresponds to the spoken utterance into the transcription:

automatically incorporate the recognized text that corresponds to the spoken utterance into the transcription; and

in response to determining to perform the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance:

cause an automated assistant to perform the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance.

13. The system of claim 12 , wherein the at least one processor is further operable to:

process, using an automatic speech recognition (ASR) model, the audio data that captures the spoken utterance to generate the recognized text that corresponds to the spoken utterance.

14. The system of claim 13 , wherein the at least one processor is further operable to:

process, using a natural language understanding (NLU) model, the recognized text that corresponds to the spoken utterance to generate annotated recognized text.

15. The system of claim 14 , wherein the at least one processor is further operable to:

determine the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, wherein determining the assistant command, that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, is based on the annotated recognized text.

16. The system of claim 12 , wherein the touch input of the user is being directed to one or more textual segments of the transcription that is displayed at the client device, wherein the touch input of the user graphically demarcates one or more of the textual segments of the transcription that is displayed at the client device, and wherein the instructions to determine whether to incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or to perform the assistant command, that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, comprise instructions to determine to perform the assistant command based on the touch input of the user graphically demarcating one or more of the textual segments of the transcription.

17. The system of claim 12 , wherein the touch input of the user is being directed to one or more fields of the transcription that is displayed at the client device, and wherein the instructions to determine whether to incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or to perform the assistant command, that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance, comprise instructions to determine to perform the assistant command based on the touch input of the user being directed to one or more fields of the transcription.

18. The system of claim 12 , wherein the instructions to automatically incorporate the recognized text that corresponds to the spoken utterance into the transcription comprise instructions to:

cause the recognized text to be visually displayed to the user via the software application accessible at the client device as part of the transcription.

19. The system of claim 18 , wherein the instructions to cause the recognized text to be visually displayed to the user via the software application accessible at the client device as part of the transcription comprise instructions to cause the recognized text to be maintained in the transcription after additional text is incorporated into the transcription.

20. A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to be operable to perform operations, the operations comprising:

receiving audio data that captures a spoken utterance of a user of a client device, the audio data being generated by one or more microphones of the client device;

determining whether touch input of the user is being simultaneously directed to a transcription, that is displayed at the client device via a software application accessible at the client device, at the same time the audio data that captures the spoken utterance is received;

in response to determining that no touch input of the user is being simultaneously directed to the transcription at the same time the audio data that captures the spoken utterance is received:

determining to incorporate recognized text, that corresponds to the spoken utterance, into the transcription;

in response to determining that touch input of the user is being simultaneously directed to the transcription at the same time the audio data that captures the spoken utterance is received:

determining, based on one or more terms of the spoken utterance, whether to:

incorporate the recognized text, that corresponds to the spoken utterance, into the transcription, or

perform an assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance;

in response to determining to incorporate the recognized text that corresponds to the spoken utterance into the transcription:

automatically incorporating the recognized text that corresponds to the spoken utterance into the transcription; and

in response to determining to perform the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance:

causing an automated assistant to perform the assistant command that is associated with the transcription and that is based on the recognized text that corresponds to the spoken utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2021
From: CARBUNE, VICTOR; ABDAGIC, ALVIN; BEHZADI, BEHSHAD; NATTA, JACOPO SANNAZZARO; PROSKURNIA, JULIA; GOJ, KRZYSZTOF ANDRZEJ; PANDIRI, SRIKANTH; ZARINS, VIESTURS; D'ERCOLE, NICOLO; SABUR, ZAHEED; KOTHARI, LUV
To: GOOGLE LLC
Reel/Frame 056340/0331 →
Continuity (1)
Related Publication 20220366910A1 · Nov 17, 2022