IP Library › Granted Patent US 12,462,797
Granted Patent B2
US 12,462,797 · App. 17/118,463 · Granted Nov 4, 2025

Rendering responses to a spoken utterance of a user utilizing a local text-response map

Inventors: Yuli Gao (Sunnyvale, CA); Sangsoo Sung (Palo Alto, CA)
Assignee: GOOGLE LLC
G10L15/22G06F3/167G10L15/26G10L15/30G06F40/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,797
App. No.
17/118,463
Granted
Nov 4, 2025
Kind
B2
Abstract

Implementations disclosed herein relate to generating and/or utilizing, by a client device, a text-response map that is stored locally on the client device. The text-response map can include a plurality of mappings, where each of the mappings define a corresponding direct relationship between corresponding text and a corresponding response. Each of the mappings is defined in the text-response map based on the corresponding text being previously generated from previous audio data captured by the client device and based on the corresponding response being previously received from a remote system in response to transmitting, to the remote system, at least one of the previous audio data and the corresponding text.

Claims (60)

1 . A method implemented by one or more processors of a client device, the method comprising:

capturing, via at least one microphone of the client device, audio data that captures a spoken utterance of a user;

processing the audio data to generate current text that corresponds to the spoken utterance, wherein processing the audio data to generate the current text utilizes a voice-to-text model stored locally on the client device;

accessing a text-response map stored locally on the client device, wherein the text-response map includes a plurality of mappings, each of the mappings defining a corresponding direct relationship between corresponding text and a corresponding response based on the corresponding text being previously generated from previous audio data captured by the client device and based on the corresponding response being previously received from a remote system in response to transmitting, to the remote system, at least one of the previous audio data and the corresponding text;

determining, by the client device, that the corresponding texts of the text-response map fail to match the current text;

in response to determining that the corresponding texts of the text-response map fail to match the current text, transmitting, to a remote system and based on determining that the corresponding texts of the text-response map fail to match the current text, the audio data or the current text;

receiving, from the remote system in response to transmitting the audio data or the current text, a response and an indication that the response is static only until an expiration event occurs, wherein the expiration event comprises determining that the user is no longer present at a particular location;

updating, in response to the indication that is received from the remote system indicating that the response is static, the text-response map by adding a given text mapping and including an indication of the expiration event with the given text mapping, the given text mapping defining a direct relationship between the current text and the response;

capturing, subsequent to updating the text-response map, second audio data;

processing the second audio data to generate a second text utilizing the voice-to-text model stored locally on the client device;

determining, based on the text-response map, that the current text matches the second text;

in response to determining that the current text matches the second text, and based on the text-response map including the given text mapping that defines the direct relationship between the current text and the response:

causing the response, from the text-response map, to be implemented; and

removing the given text mapping from the text-response map when the expiration event occurs.

2 . The method of claim 1 , wherein the spoken utterance comprise a query or command and the response comprises a response to the query or command.

3 . The method of claim 2 , wherein the response comprises information indicated to be of interest in the query or command.

4 . The method of claim 1 , wherein updating the text-response map includes removing one or more mappings from the text-response map.

5 . The method of claim 1 , wherein causing the response to be implemented comprises causing the client device to render the response.

6 . The method of claim 1 , wherein the response comprises a command and wherein causing the response to be implemented comprises causing the command to be provided to a smart device to alter a state of the smart device.

7 . A client device comprising:

one or more processors;

microphones; and

memory storing computer-executable instructions which, when executed by the one or more processors, causes the one or more processors to:

capture, via at least one of the microphones, audio data that captures a spoken utterance of a user;

process the audio data to generate current text that corresponds to the spoken utterance, wherein processing the audio data to generate the current text utilizes a voice-to-text model stored locally on the client device;

access a text-response map stored locally on the client device, wherein the text-response map includes a plurality of mappings, each of the mappings defining a corresponding direct relationship between corresponding text and a corresponding response based on the corresponding text being previously generated from previous audio data captured by the client device and based on the corresponding response being previously received from a remote system in response to transmitting, to the remote system, at least one of the previous audio data and the corresponding text;

determine, by the client device, that the corresponding texts of the text-response map fail to match the current text;

in response to determining that the corresponding texts of the text-response map fail to match the current text, transmit, to a remote system and based on determining that the corresponding texts of the text-response map fail to match the current text, the audio data or the current text;

receive, from the remote system in response to transmitting the audio data or the current text, a response and an indication that the response is static only until an expiration event occurs, wherein the expiration event comprises determining that the user is no longer present at a particular location;

update, in response to the indication that is received from the remote system indicating that the response is static, the text-response map by adding a given text mapping and including an indication of the expiration event with the given text mapping, the given text mapping defining a direct relationship between the current text and the response;

capture, subsequent to updating the text-response map, second audio data;

process the second audio data to generate a second text utilizing the voice-to-text model stored locally on the client device;

determine, based on the text-response map, that the current text matches the second text;

in response to determining that the current text matches the second text, and based on the text-response map including the given text mapping that defines the direct relationship between the current text and the response:

cause the response, from the text-response map, to be implemented; and

remove the given text mapping from the text-response map when the expiration event occurs.

8 . The client device of claim 7 , wherein the spoken utterance comprises a query or command and the response comprises a response to the query or command.

9 . The client device of claim 8 , wherein the response comprises information indicated to be of interest in the query or command.

10 . The client device of claim 7 , wherein in updating the text-response map, one or more of the processors are to remove one or more mappings from the text- response map.

11 . The client device of claim 7 , wherein in causing the response to be implemented, one or more of the processors are to cause the client device to render the response.

12 . The client device of claim 7 , wherein the response comprises a command and wherein in causing the response to be implemented, one or more of the processors are to cause the command to be provided to a smart device to alter a state of the smart device.

13 . A non-transitory computer readable storage medium configured to store instructions that, when executed by one or more processors, cause one or more of the processors to:

capture, via at least one microphone of a client device, audio data that captures a spoken utterance of a user;

process the audio data to generate current text that corresponds to the spoken utterance, wherein processing the audio data to generate the current text utilizes a voice-to-text model stored locally on the client device;

access a text-response map stored locally on the client device, wherein the text-response map includes a plurality of mappings, each of the mappings defining a corresponding direct relationship between corresponding text and a corresponding response based on the corresponding text being previously generated from previous audio data captured by the client device and based on the corresponding response being previously received from a remote system in response to transmitting, to the remote system, at least one of the previous audio data and the corresponding text;

determine, by the client device, that the corresponding texts of the text-response map fail to match the current text;

in response to determining that the corresponding texts of the text-response map fail to match the current text, transmit, to a remote system and based on determining that the corresponding texts of the text-response map fail to match the current text, the audio data or the current text;

receive, from the remote system in response to transmitting the audio data or the current text, a response and an indication that the response is static only until an expiration event occurs, wherein the expiration event comprises determining that the user is no longer present at a particular location;

update, in response to the indication that is received from the remote system indicating that the response is static, the text-response map by adding a given text mapping and including an indication of the expiration event with the given text mapping, the given text mapping defining a direct relationship between the current text and the response;

capture, subsequent to updating the text-response map, second audio data;

process the second audio data to generate a second text utilizing the voice-to-text model stored locally on the client device;

determine, based on the text-response map, that the current text matches the second text;

in response to determining that the current text matches the second text, and based on the text-response map including the given text mapping that defines the direct relationship between the current text and the response:

cause the response, from the text-response map, to be implemented; and

remove the given text mapping from the text-response map when the expiration event occurs.

14 . The non-transitory computer readable storage medium of claim 13 , wherein the spoken utterance comprises a query or command and the response comprises a response to the query or command.

15 . The non-transitory computer readable storage medium of claim 14 , wherein the response comprises information indicated to be of interest in the query or command.

16 . The non-transitory computer readable storage medium of claim 13 , wherein in updating the text-response map, one or more of the processors are to remove one or more mappings from the text-response map.

17 . The non-transitory computer readable storage medium of claim 13 , wherein in causing the response to be implemented, one or more of the processors are to cause the client device to render the response.

18 . The non-transitory computer readable storage medium of claim 13 , wherein the response comprises a command and wherein in causing the response to be implemented, one or more of the processors are to cause the command to be provided to a smart device to alter a state of the smart device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: GAO, YULI; SUNG, SANGSOO
To: GOOGLE LLC
Reel/Frame 058720/0004 →
Continuity (2)
Continuation 16609403
Related Publication 20210097999A1 · Apr 1, 2021
References Cited (120)
US 6295535B1 · Radcliffe · 2001 [cited by examiner]
US 7240094B2 · Hackney · 2007 [cited by examiner]
US 8117531B1 · Lueck · 2012 [cited by examiner]
US 8311835B2 · Lecoeuche · 2012 [cited by examiner]
US 8868409B1 · Mengibar et al. · 2014 [cited by applicant]
US 8972180B1 · Zhao · 2015 [cited by examiner]
US 9124472B1 · Schneider · 2015 [cited by examiner]
US 9300761B2 · Ryu · 2016 [cited by examiner]
US 9679568B1 · Taubman · 2017 [cited by examiner]
US 10235999B1 · Naughton et al. · 2019 [cited by applicant]
US 10388277B1 · Ghosh et al. · 2019 [cited by applicant]
US 10777203B1 · Pasko · 2020 [cited by examiner]
US 10891958B2 · Gao et al. · 2021 [cited by applicant]
US 20070179789A1 · Bennett · 2007 [cited by applicant]
US 20080104043A1 · Garg · 2008 [cited by examiner]
US 20080115141A1 · Welingkar · 2008 [cited by examiner]
US 20080244556A1 · Plante et al. · 2008 [cited by applicant]
US 20090192968A1 · Tunstall-Pedoe · 2009 [cited by examiner]
US 20100005081A1 · Bennett · 2010 [cited by applicant]
US 20110106617A1 · Cooper · 2011 [cited by examiner]
US 20110106966A1 · Smit · 2011 [cited by examiner]
US 20110307435A1 · Overell · 2011 [cited by examiner]
US 20120179469A1 · Newman · 2012 [cited by applicant]
US 20120265531A1 · Bennett · 2012 [cited by examiner]
US 20130006626A1 · Aiyer · 2013 [cited by examiner]
US 20130326384A1 · Moore · 2013 [cited by examiner]
US 20130326407A1 · van Os · 2013 [cited by examiner]
US 20140122059A1 · Patel et al. · 2014 [cited by applicant]
US 20140195230A1 · Han · 2014 [cited by examiner]
US 20140250195A1 · Capper et al. · 2014 [cited by applicant]
US 20140280169A1 · Liu · 2014 [cited by examiner]
US 20140370920A1 · Caillette · 2014 [cited by examiner]
US 20150058488A1 · Backholm · 2015 [cited by applicant]
US 20150084770A1 · Xiao · 2015 [cited by examiner]
US 20150095267A1 · Behere · 2015 [cited by examiner]
US 20150170653A1 · Berndt et al. · 2015 [cited by applicant]
US 20150187232A1 · Bailiang · 2015 [cited by examiner]
US 20150248464A1 · Desai · 2015 [cited by examiner]
US 20150279352A1 · Willett et al. · 2015 [cited by applicant]
US 20150310755A1 · Haverlock et al. · 2015 [cited by applicant]
US 20150341486A1 · Knighton · 2015 [cited by applicant]
US 20160110415A1 · Clark · 2016 [cited by examiner]
US 20160262017A1 · Lavee et al. · 2016 [cited by applicant]
US 20160275075A1 · Yu · 2016 [cited by examiner]
US 20160292204A1 · Klemm et al. · 2016 [cited by applicant]
US 20160350320A1 · Sung et al. · 2016 [cited by applicant]
US 20170004204A1 · Bastide et al. · 2017 [cited by applicant]
US 20170083969A1 · Takeda · 2017 [cited by examiner]
US 20170192976A1 · Bhatia et al. · 2017 [cited by applicant]
US 20170235825A1 · Gordon · 2017 [cited by applicant]
US 20170236519A1 · Jung et al. · 2017 [cited by applicant]
US 20180092189A1 · Reier · 2018 [cited by examiner]
US 20180122366A1 · Nishikawa · 2018 [cited by applicant]
US 20180131645A1 · Magliozzi · 2018 [cited by examiner]
US 20180150739A1 · Wu · 2018 [cited by applicant]
US 20180255180A1 · Goldberg et al. · 2018 [cited by applicant]
US 20180301148A1 · Roman et al. · 2018 [cited by applicant]
US 20180341643A1 · Alders · 2018 [cited by examiner]
US 20190027130A1 · Tsunoo · 2019 [cited by examiner]
US 20190066696A1 · Mu et al. · 2019 [cited by applicant]
US 20190129688A1 · Yao · 2019 [cited by examiner]
US 20190129938A1 · Yao · 2019 [cited by examiner]
US 20190173913A1 · Kras · 2019 [cited by applicant]
US 20190173914A1 · Irimie et al. · 2019 [cited by applicant]
US 20190222555A1 · Skinner et al. · 2019 [cited by applicant]
US 20190279620A1 · Talwar et al. · 2019 [cited by applicant]
US 20190295552A1 · Pasko · 2019 [cited by examiner]
US 20190310804A1 · Fukuda · 2019 [cited by examiner]
US 20190312973A1 · Engelke et al. · 2019 [cited by applicant]
US 20190354630A1 · Guo · 2019 [cited by examiner]
US 20190371314A1 · Naughton et al. · 2019 [cited by applicant]
US 20190392037A1 · Guo et al. · 2019 [cited by applicant]
US 20200218767A1 · Ritchey et al. · 2020 [cited by applicant]
US 20200220935A1 · Bao · 2020 [cited by examiner]
US 20210097236A1 · Fujimoto · 2021 [cited by examiner]
CN 1735929A · 2006 [cited by examiner]
CN 102629246 · 2012 [cited by applicant]
CN 103247291 · 2013 [cited by applicant]
CN 103295575 · 2013 [cited by applicant]
CN 105261366 · 2016 [cited by applicant]
CN 105793923 · 2016 [cited by applicant]
CN 106796496 · 2017 [cited by applicant]
CN 107660303 · 2018 [cited by applicant]
CN 108010523 · 2018 [cited by applicant]
CN 108170034 · 2018 [cited by applicant]
JP 3960188B2 · 2007 [cited by examiner]
JP 2007527640 · 2007 [cited by applicant]
JP 2013205524 · 2013 [cited by applicant]
JP 2013207726A · 2013 [cited by examiner]
JP 2014071449 · 2014 [cited by applicant]
JP 2015015675A · 2015 [cited by examiner]
JP 2015106203 · 2015 [cited by applicant]
JP 2017072725 · 2017 [cited by applicant]
JP 2012501480 · 2018 [cited by applicant]
KR 20060070605 · 2006 [cited by applicant]
KR 20130086971 · 2013 [cited by applicant]
WO 2005006116 · 2005 [cited by applicant]
WO WO2012167276A1 · 2012 [cited by examiner]
WO 2014051207 · 2014 [cited by applicant]
WO 2016122902 · 2016 [cited by applicant]
WO 2017146803 · 2017 [cited by applicant]
WO WO2018016669A2 · 2018 [cited by examiner]
WO 2017141502 · 2018 [cited by applicant]
WO 2020005241 · 2020 [cited by applicant]
Japanese Patent Office; Notice of Allowance issued in Application No. 2020-569950; 3 pages; dated Jun. 21, 2021. [cited by applicant]
European Patent Office; Communication Issued in Application No. 20195601.8; 10 pages; dated Feb. 17, 2021. [cited by applicant]
Intellectual Property India; Examination Report issued in Application No. 202027052364; 7 pages; dated Dec. 10, 2022. [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 18743935.1; 46 pages; dated Feb. 18, 2020. [cited by applicant]
European Patent Office; Commincation under Rule 71(3) EPC issued in Application No. 18743935.1; 46 pages; dated Apr. 30, 2020. [cited by applicant]
International Search Report and Written Opinion issued is Application No. PCT/US2018/039850 dated Feb. 25, 2019. [cited by applicant]
Japanese Patent Office; Notice of Reasons for Rejection issued in app. No. 2021-118613, 11 pages, dated Aug. 22, 2022. [cited by applicant]
European Patent Office; Communication pursuant to Article 94(3) issued in Application No. 20195601.8, 5 pages, dated Aug. 24, 2022. [cited by applicant]
Korean Intellectual Property Office; Notice of Office Action issued in Application Ser. No. KR10-2020-7037198; 13 pages; dated Aug. 18, 2022. [cited by applicant]
Korean Intellectual Property Office; Notice of Allowance issued for Application No. 10-2020-7037198, 4 pages, dated Jan. 2, 2023. [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 20195601.8, 73 pages, dated Apr. 26, 2023. [cited by applicant]
Korean Intellectual Property Office; Notice of Office Action issued in Application Ser. No. KR10-2023-7010851; 6 pages; dated May 25, 2023. [cited by applicant]
Korean Intellectual Property Office: Notice of Allowance issued for Application No. 10-2023-7010851, 5 pages, dated Sep. 5, 2023. [cited by applicant]
China National Intellectual Property Administration; Notice of Grant issued in Application No. 201880094598.1; 4 pages; dated Apr. 25, 2024. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 201880094598.1; 31 pages; dated Dec. 1, 2023. [cited by applicant]
Intellectual Property India; Hearing Notice issued for Application No. 202027052364, 2 pages, dated Mar. 11, 2024. [cited by applicant]