IP Library Granted Patent US 11,314,481
Granted Patent B2
US 11,314,481 · App. 15/781,787 · Granted Apr 26, 2022

Systems and methods for voice-based initiation of custom device actions

Inventors: Bo Wang (San Jose, CA); Venkat Kotla (Mountain View, CA); Chad Yoshikawa (Mountain View, CA); Chris Ramsdale (Mountain View, CA); Pravir Gupta (Los Altos, CA); Alfonso Gomez-Jordana (Mountain View, CA); Kevin Yeun (Mountain View, CA); Jae Won Seo (Mountain View, CA); Lantian Zheng (San Jose, CA); Sang Soo Sung (Palo Alto, CA)
Assignee: GOOGLE LLC
G06F3/167G10L15/1822G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,314,481
App. No.
15/781,787
Granted
Apr 26, 2022
Kind
B2
Abstract

Systems and methods for enabling voice-based interactions with electronic devices can include a data processing system maintaining a plurality of device action data sets and a respective identifier for each device action data set. The data processing system can receive, from an electronic device, an audio signal representing a voice query and an identifier. The data processing system can identify, using the identifier, a device action data set. The data processing system can identify a device action from device action data set based on content of the audio signal. The data processing system can then identify, from the device action dataset, a command associated with the device action and send the command to the for execution device for execution.

Claims (48)

1. A data processing system to enable voice-based interactions with client devices, comprising:

a communications interface to receive, from a computing device, device action data, a device model identifier defining a device model of a plurality of client devices that are associated with the device action data, and a plurality of audio or visual responses for presentation by the plurality of client devices associated with the device model, the device action data indicative of a plurality of device actions that are supported by the plurality of client devices associated with the device model, each of the plurality of device actions being associated with a corresponding device executable command of a plurality of device executable commands that are supported by the plurality of client devices associated with the device model, each device executable command of the plurality of device executable commands being an executable command specific to the plurality of client devices associated with the device model to trigger execution of a corresponding device action of the plurality of device actions, and each audio or visual response of the plurality of audio or visual responses being for presentation in connection with performance of a corresponding device action of the plurality of device actions supported by the plurality of client devices associated with the device model;

a memory to store the device action data;

a device action customization component to map the device action data to the device model identifier and to map each of the plurality of device actions to the corresponding audio or visual response for presentation in connection with the device action;

the communications interface to receive, from a client device of the plurality of client devices associated with the device model, an audio signal and the device model identifier, the audio signal obtained by the client device responsive to a voice-based query;

a natural language processor component to identify, using the device model identifier and content associated with the audio signal, a device action of the plurality of device actions supported by the plurality of client devices associated with the device model;

the device action customization component to identify a device executable command of the plurality of device executable commands corresponding to the device action and to identify an audio or visual response of the plurality of audio or visual responses corresponding to the device action; and

the communications interface to transmit, to the client device, the device executable command for execution responsive to the voice-based query to cause performance of the device action and to transmit, to the client device, the audio or visual response for rendering on the client device in connection with and based on the device action.

2. The data processing system of claim 1 comprising:

a speech recognition component to convert the audio signal received from the client device into a corresponding text, the natural language processor component identifying the device action using the corresponding text.

3. The data processing system of claim 1 , comprising the natural language processor component to:

determine, for each first device action of the plurality of device actions, a corresponding weight value for matching the content of the received audio signal with the first device action; and

identify the device action based on the weight values.

4. The data processing system of claim 1 , comprising:

the device action customization component to validate the device model prior to interacting with the client devices associated with the device model and upon successful testing of each of the plurality of device executable commands.

5. The data processing system of claim 1 comprising:

the device action customization component to provide a user interface to allow the computing device to provide the device model identifier and the device action data.

6. The data processing system of claim 1 comprising:

the device action customization component to provide a restful application programming interface (API) to allow transmission of the device model identifier and the device action data to the data processing system.

7. The data processing system of claim 1 , wherein the device model identifier is associated with an application installed on the plurality of client devices, the plurality of device actions are device actions supported by the application, and the plurality of executable commands are executable commands specific to the application.

8. The data processing system of claim 7 , comprising:

the device action customization component to identify one or more parameters associated with the device executable command; and

the communications interface to transmit, to the client device, the one or more parameters.

9. A method of enabling voice-based interactions with client devices, the method comprising:

receiving, from a computing device, device action data, a device model identifier defining a device model of a plurality of client devices that are associated with the device action data, and a plurality of audio or visual responses for presentation by the plurality of client devices associated with the device model, the device action data indicative of a plurality of device actions that are supported by the plurality of client devices associated with the device model, each of the plurality of device actions being associated with a corresponding device executable command of a plurality of device executable commands that are supported by the plurality of client devices associated with the device model, each device executable command of the plurality of device executable commands being an executable command specific to the plurality of client devices associated with the device model to trigger execution of a corresponding device action of the plurality of device actions, and each audio or visual response of the plurality of audio or visual responses being for presentation in connection with performance of a corresponding device action of the plurality of device actions supported by the plurality of client devices associated with the device model;

storing the device action data;

mapping the device action data to the device model identifier and mapping each of the plurality of device actions to the corresponding audio or visual response for presentation in connection with the device action;

receiving, from a client device of the plurality of client devices associated with the device model, an audio signal and the device model identifier, the audio signal obtained by the client device responsive to a voice-based query;

identifying, using the device model identifier and content associated with the audio signal, a device action of the plurality of device actions supported by the plurality of client devices associated with the device model;

identifying a device executable command of the plurality of device executable commands corresponding to the device action, and identifying an audio or visual response of the plurality of audio or visual responses corresponding to the device action; and

transmitting, to the client device, the device executable command for execution responsive to the voice-based query to cause performance of the device action, and transmitting, to the client device, the audio or visual response for rendering on the client device in connection with and based on the device action.

10. The method of claim 9 comprising:

converting the audio signal received from the client device into a corresponding text, a natural language processor component identifying the device action using the corresponding text.

11. The method of claim 9 , comprising:

determining, for each first device action of the plurality of device actions, a corresponding weight value for matching the content of the received audio signal with the first device action; and

identifying the device action based on the weight values.

12. The method of claim 9 , comprising:

validating the device model prior to interacting with the client devices associated with the device model and upon successful testing of each of the plurality of device executable commands.

13. The method of claim 9 comprising:

providing a user interface to allow the computing device to provide the device model identifier and the device action data.

14. The method of claim 9 comprising:

the device action customization component to provide a restful application programming interface (API) for use by the computing device to transmit the device model identifier and the device action data.

15. The method of claim 9 , wherein the transmitting, to the client device, the device executable command comprises transmitting the device executable command in a JSON response provided to the client device.

16. The method of claim 9 , wherein the transmitting, to the client device, the audio or visual response comprises transmitting, to the client device, the audio or visual response to be rendered by the client device by converting a text expression to audio.

17. The method of claim 9 , wherein the device model identifier is associated with an application installed on the plurality of client devices, the plurality of device actions are device actions supported by the application, and the plurality of executable commands are executable commands specific to the application.

18. The method of claim 17 comprising:

identifying one or more parameters associated with the device executable command; and

transmitting, to the client device, the one or more parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2019
From: WANG, BO; VENKATA, SUBBAIAH; YOSHIKAWA, CHAD; RAMSDALE, CHRIS; GUPTA, PRAVIR; GOMEZ-JORDANA, ALFONSO; YEUN, KEVIN; SEO, JAE WON; ZHENG, LANTIAN; SUNG, SANG SOO
To: GOOGLE LLC
Reel/Frame 050477/0198 →
Continuity (2)
Provisional Application 62640007 · Mar 7, 2018
Related Publication 20210026593A1 · Jan 28, 2021
Cited By (1)
US 12,548,566