IP Library › Granted Patent US 12,327,559
Granted Patent B2
US 12,327,559 · App. 18/430,253 · Granted Jun 10, 2025

Multimodal responses

Inventors: April Pufahl (Mountain View, CA); Jared Strawderman (San Jose, CA); Harry Yu (San Francisco, CA); Adriana Olmos Antillon (San Francisco, CA); Jonathan Livni (San Francisco, CA); Okan Kolak (Sunnyvale, CA); James Giangola (Mountain View, CA); Nitin Khandelwal (Sunnyvale, CA); Jason Kearns (Oakland, CA); Andrew Watson (Zurich, CH); Joseph Ashear (Redwood City, CA); Valerie Nygaard (Saratoga, CA)
Assignee: GOOGLE LLC
G10L15/22G06F1/1694G06F3/167G06F2203/0381G10L2015/223G10L2015/225H04M2203/251H04M2203/253
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,559
App. No.
18/430,253
Filed
Feb 1, 2024
Granted
Jun 10, 2025
Kind
B2
Art Unit
2654
USPC
704/275
Abstract

Systems, methods, and apparatus for using a multimodal response in the dynamic generation of client device output that is tailored to a current modality of a client device is disclosed herein. Multimodal client devices can engage in a variety of interactions across the multimodal spectrum including voice only interactions, voice forward interactions, multimodal interactions, visual forward interactions, visual only interactions etc. A multimodal response can include a core message to be rendered for all interaction types as well as one or more modality dependent components to provide a user with additional information.

Claims (56)

1. A method implemented by one or more processors, the method comprising:

determining a client device action based on one or more instances of user interface input provided by a user of a multimodal client device;

processing, using a machine learning model, (1) sensor data from one or more sensors of the multimodal client device and (2) a multimodal response, to generate client device output,

wherein the sensor data from the one or more sensors of the multimodal client device indicate a current client device modality of the multimodal client device,

wherein the current client device modality is one of a plurality of discrete client device modalities available for the multimodal client device,

wherein the sensor data is in addition to any sensor data generated by the one or more instances of user interface input,

wherein the multimodal response includes components of output for the client device action for the plurality of discrete client device modalities,

wherein the multimodal response is received by the multimodal client device from a remote server,

wherein generating the client device output is done by the multimodal client device, and

wherein the components of output include at least (a) a core message component representing information to render for each of the plurality of discrete client device modalities and (b) one or more modality dependent components each representing corresponding information to render for one or more of the plurality of discrete client device modalities;

causing the client device output to be rendered by one or more user interface output devices of the multimodal client device;

while at least part of the client device output is being rendered by the one or more user interface output devices of the multimodal client device:

detecting a switch of the multimodal client device, based on alternative sensor data from the one or more sensors of the multimodal client device, from the current client device modality to a discrete new client device modality;

in response to detecting the switch, generating alternative client device output based on processing (1) the alternative sensor data from the one or more sensors of the multimodal client device and (2) the multimodal response, using the machine learning model; and

causing the alternative client device output to be rendered by the one or more user interface output devices of the multimodal client device.

2. The method of claim 1 , wherein the multimodal response is received by the multimodal client device from the remote server in response to a request, transmitted to the remote server by the client device, that is based on the user interface input, and wherein determining the current client device modality of the multimodal client device is by the multimodal client device and occurs after transmission of the request.

3. The method of claim 1 ,

wherein the client device output includes audible output rendered via at least one speaker of the one or more user interface output devices of the multimodal client device and visual output rendered via at least one display of the one or more user interface output devices,

wherein the alternative client device output lacks the visual output, and

wherein causing the alternative client device output to be rendered by the multimodal client device comprises ceasing rendering of the visual output by the at least one display.

4. The method of claim 1 , wherein the current client device modality is a voice only interaction and the client device output is rendered via one or more speakers of the one or more user interface output devices.

5. The method of claim 1 , wherein the current client device modality is a voice forward interaction, the core message component of the client device output is rendered via only one or more speakers of the one or more user interface output devices, and the one or more modality dependent components of the client device output are rendered via a touch screen of the one or more user interface output devices.

6. The method of claim 1 , wherein the current client device modality is a multimodal interaction, the client device output is rendered via one or more speakers and via a touch screen of the one or more user interface output devices.

7. The method of claim 1 , wherein the current device modality is a visual forward interaction, the core message component of the client device output is rendered via only a touch screen of the one or more user interface output devices, and the one or more modality dependent components of the client device output are rendered via one or more speakers of the one or more user interface output devices.

8. The method of claim 1 , wherein the current device modality is a visual only interaction, and the client device output is rendered via only a touch screen of the one or more user interface output devices.

9. A multimodal client device, comprising:

one or more microphones;

one or more sensors that are in addition to the one or more microphones;

one or more speakers;

one or more displays;

memory storing instructions;

one or more processors executing the instructions to perform the method of:

determining a client device action based on one or more instances of user interface input provided by a user of the multimodal client device;

processing, using a machine learning model, (1) sensor data from the one or more sensors of the multimodal client device and (2) a multimodal response, to generate client device output,

wherein the sensor data from the one or more sensors of the multimodal client device indicate a current client device modality of the multimodal client device,

wherein the current client device modality is one of a plurality of discrete client device modalities available for the multimodal client device,

wherein the sensor data is in addition to any sensor data generated by the one or more instances of user interface input,

wherein the multimodal response includes components of output for the client device action for the plurality of discrete client device modalities,

wherein the multimodal response is received by the multimodal client device from a remote server,

wherein generating the client device output is done by the multimodal client device, and

wherein the components of output include at least (a) a core message component representing information to render for each of the plurality of discrete client device modalities and (b) one or more modality dependent components each representing corresponding information to render for one or more of the plurality of discrete client device modalities;

causing the client device output to be rendered by one or more user interface output devices of the multimodal client device;

while at least part of the client device output is being rendered by the one or more user interface output devices of the multimodal client device:

detecting a switch of the multimodal client device, based on alternative sensor data from the one or more sensors of the multimodal client device, from the current client device modality to a discrete new client device modality;

in response to detecting the switch, generating alternative client device output based on processing (1) the alternative sensor data from the one or more sensors of the multimodal client device and (2) the multimodal response, using the machine learning model; and

causing the alternative client device output to be rendered by the one or more user interface output devices of the multimodal client device.

10. The multimodal client device of claim 9 , wherein the multimodal response is received by the multimodal client device from the remote server in response to a request, transmitted to the remote server by the client device, that is based on the user interface input, and wherein determining the current client device modality of the multimodal client device is by the multimodal client device and occurs after transmission of the request.

11. The multimodal client device of claim 9 ,

wherein the client device output includes audible output rendered via at least one speaker of the one or more user interface output devices of the multimodal client device and visual output rendered via at least one display of the one or more user interface output devices,

wherein the alternative client device output lacks the visual output, and

wherein causing the alternative client device output to be rendered by the multimodal client device comprises ceasing rendering of the visual output by the at least one display.

12. The multimodal client device of claim 9 , wherein the current client device modality is a voice only interaction and the client device output is rendered via one or more speakers of the one or more user interface output devices.

13. The multimodal client device of claim 9 , wherein the current client device modality is a voice forward interaction, the core message component of the client device output is rendered via only one or more speakers of the one or more user interface output devices, and the one or more modality dependent components of the client device output are rendered via a touch screen of the one or more user interface output devices.

14. The multimodal client device of claim 9 , wherein the current client device modality is a multimodal interaction, the client device output is rendered via one or more speakers and via a touch screen of the one or more user interface output devices.

15. The multimodal client device of claim 9 , wherein the current device modality is a visual forward interaction, the core message component of the client device output is rendered via only a touch screen of the one or more user interface output devices, and the one or more modality dependent components of the client device output are rendered via one or more speakers of the one or more user interface output devices.

16. The multimodal client device of claim 9 , wherein the current device modality is a visual only interaction, and the client device output is rendered via only a touch screen of the one or more user interface output devices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2024
From: PUFAHL, APRIL; STRAWDERMAN, JARED; YU, HARRY; ANTILLON, ADRIANA OLMOS; LIVINI, JONATHAN; KOLAK, OKAN; GIANGOLA, JAMES; KHANDELWAL, NITIN; KEAMS, JAMES; `WATSON, ANDREW; ASHEAR, JOSEPH; NYGAARD, VALERIE
To: GOOGLE LLC
Reel/Frame 066771/0875 →
Continuity (4)
Continuation 17515901 · Nov 1, 2021
Continuation 16251982 · Jan 18, 2019
Provisional Application 62726947 · Sep 4, 2018
Related Publication 20240169989A1 · May 23, 2024
References Cited (52)
US 8370160B2 · Pearce · 2013 [cited by applicant]
US 8386260B2 · Engelsma · 2013 [cited by applicant]
US 8977965B1 · Ehlen et al. · 2015 [cited by applicant]
US 9443519B1 · Bakshi et al. · 2016 [cited by applicant]
US 9507439B2 · Yoon · 2016 [cited by applicant]
US 9690542B2 · Reddy · 2017 [cited by examiner]
US 10031549B2 · Costa · 2018 [cited by applicant]
US 10043516B2 · Saddler et al. · 2018 [cited by applicant]
US 10134397B2 · Bakshi · 2018 [cited by applicant]
US 10671428B2 · Zeitlin · 2020 [cited by applicant]
US 11200893B2 · Kirazci · 2021 [cited by examiner]
US 11217240B2 · Huber · 2022 [cited by applicant]
US 20030167167A1 · Gong · 2003 [cited by examiner]
US 20060072542A1 · Sinnreich et al. · 2006 [cited by applicant]
US 20090182562A1 · Caire · 2009 [cited by applicant]
US 20120109868A1 · Murillo · 2012 [cited by examiner]
US 20120198339A1 · Williams et al. · 2012 [cited by applicant]
US 20130235073A1 · Jaramillo · 2013 [cited by examiner]
US 20160179908A1 · Johnston et al. · 2016 [cited by applicant]
US 20160198988A1 · Bhavaraju et al. · 2016 [cited by applicant]
US 20160299959A1 · Sankar et al. · 2016 [cited by applicant]
US 20170289766A1 · Scott · 2017 [cited by examiner]
US 20170329573A1 · Mixter · 2017 [cited by applicant]
US 20180329677A1 · Gruber et al. · 2018 [cited by applicant]
US 20200075002A1 · Pufahl et al. · 2020 [cited by applicant]
US 20200098368A1 · Lemay et al. · 2020 [cited by applicant]
US 20220051675A1 · Pufahl et al. · 2022 [cited by applicant]
US 20230026521A1 · Gray · 2023 [cited by examiner]
US 20230343330A1 · Machanavajhala · 2023 [cited by examiner]
CN 101689187 · 2010 [cited by applicant]
CN 101911064 · 2010 [cited by applicant]
CN 102340649 · 2012 [cited by applicant]
CN 102818913 · 2012 [cited by applicant]
CN 203135171 · 2013 [cited by applicant]
CN 104423829 · 2015 [cited by applicant]
CN 104618206 · 2015 [cited by applicant]
CN 105721666 · 2016 [cited by applicant]
CN 106025679 · 2016 [cited by applicant]
CN 107423809 · 2017 [cited by applicant]
CN 108062213 · 2018 [cited by applicant]
CN 108106541 · 2018 [cited by applicant]
CN 108197329 · 2018 [cited by applicant]
WO 2012103321 · 2012 [cited by applicant]
WO 2017197010 · 2017 [cited by applicant]
Chinese patent office office action dated Jun. 23, 2023 for CN201910826487. (Year: 2023). [cited by examiner]
Translation of Chinese patent office office action dated Jun. 23, 2023 for CN201910826487. (Year: 2023). [cited by examiner]
China National Intellectual Property Administration; Notice of Grant issued in Application No. 201910826487.1; 4 pages; dated Nov. 20, 2023. [cited by applicant]
China National Intellectual Property Administration; Notification of Second Office Action issued in Application No. 201910826487.1; 17 pages; dated Jun. 30, 2023. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 201910826487.1; 27 pages; dated Dec. 5, 2022. [cited by applicant]
Google Developers; Finding the Right Voice Interactions for Your App (Google I/O '17) https:/www.youtube.com/watch?v=0PmWruLLUoE&t=4s May 18, 2017. [cited by applicant]
Google Developers; Defining Multimodal Interactions One Size Does Not Fit All (Google I/O '17) https://www.youtube.com/watch?v=fw27RFHP2tc&t=1s May 18, 2017. [cited by applicant]
Google Developers; Design Actions for the Google Assistant beyond smart speakers (Google I/O '18) https://www.youtube.com/watch?v=JDakZMIXpQo&t=623s May 9, 2018. [cited by applicant]