Systems and methods for enhanced contextual responses with a virtual assistant
A voice query is received via an input to a virtual assistant from a user. The virtual assistant determines a media context from first media content, the media context being colocated and contemporaneous with the voice query. The voice query is processed to generate a textual query and to identify a keyword from the voice query, and a response content provider is selected based on the keyword and the media context. The textual query and the media context are communicated to the response content provider via a network interface. Query response data is received from the response content provider via the network interface, the query response data comprising voice data. The virtual assistant then generates, at an output, a query response based on the query response data.
1 . A method, performed by a first device of a plurality of devices, comprising:
receiving a voice query;
determining a media context from first media content, the media context being colocated and contemporaneous with the voice query;
processing the voice query to generate a textual query;
performing a keyword identification skill to identify a keyword from the voice query;
determining a character depicted in the first media content based, at least in part, on the textual query;
identifying a selected response content provider from a plurality of response content providers based on the keyword, the media context, and the determined character, wherein each of the plurality of response content providers own respective copyrights to respective characters;
communicating the textual query, the media context, and the determined character to the selected response content provider via a network interface;
receiving query response data from the selected response content provider via the network interface, wherein the query response data comprises voice data generated based on the keyword, the media context, and the determined character, and wherein the voice data simulates a voice of the determined character;
generating, at an output of the first device, a query response based on the query response data;
receiving, at the first device, an indication to configure a second device of the plurality of devices to perform the keyword identification skill; and
sharing, with the second device, the keyword identification skill, wherein sharing the keyword identification skill with the second device causes the second device to be configured to identify an association between the keyword and the selected response content provider.
2 . The method of claim 1 , wherein determining the media context comprises determining metadata from the media context, and wherein communicating the textual query and the media context to the selected response content provider includes communicating the metadata to the selected response content provider.
3 . The method of claim 1 , wherein the query response comprises response media content generated based on the query response data, the response media content comprising audio content based on the voice data.
4 . The method of claim 3 , wherein the query response data comprises character rendering data based on the media context, and wherein the response media content is generated based on the character rendering data.
5 . The method of claim 3 , wherein the query response data comprises scenery rendering data for scenery based on the media context, and wherein the response media content is generated based on the scenery rendering data.
6 . The method of claim 3 , wherein the response media content comprises a visual representation of the determined character based on the media context.
7 . The method of claim 3 , wherein the response media content comprises visual scenery based on the media context.
8 . The method of claim 1 , wherein the query response data comprises response media content generated by the selected response content provider.
9 . The method of claim 8 , wherein the response media content comprises a visual representation of the determined character based on the media context and the query response data.
10 . The method of claim 8 , wherein the response media content comprises visual scenery based on the media context and the query response data.
11 . The method of claim 1 , wherein processing the voice query to identify the keyword from the voice query comprises:
transmitting the textual query to a search engine via the network interface;
receiving a search result from the search engine via the network interface; and
identifying the keyword from the search result.
12 . A system comprising:
input/output (I/O) circuitry configured to receive a voice query;
a network interface; and
control circuitry of a first device of a plurality of devices configured to:
determine a media context from first media content, the media context being colocated and contemporaneous with the voice query;
process the voice query to generate a textual query;
perform a keyword identification skill to identify a keyword from the voice query;
determine a character depicted in the first media content based, at least in part, on the textual query;
identify a selected response content provider from a plurality of response content providers based on the keyword, the media context, and the determined character, wherein each of the plurality of response content providers own respective copyrights to respective characters; and
communicate the textual query, the media context, and the determined character to the selected response content provider via the network interface;
wherein the I/O circuitry is further configured to:
receive query response data from the selected response content provider via the network interface, wherein the query response data comprises voice data generated based on the keyword, the media context, and the determined character, and wherein the voice data simulates a voice of the determined character;
generate, at an output of the first device, a query response based on the query response data; and
receive, at the first device, an indication to configure a second device of the plurality of devices to perform the keyword identification skill; and
wherein the control circuitry is further configured to share, with the second device, the keyword identification skill, wherein sharing the keyword identification skill with the second device causes the second device to be configured to identify an association between the keyword and the selected response content provider.
13 . The system of claim 12 , wherein the control circuitry is further configured to determine metadata from the media context and to communicate the metadata to the selected response content provider with the textual query and the media context.
14 . The system of claim 12 , wherein the query response comprises response media content, and wherein the control circuitry is further configured to generate the response media content based on the query response data and the voice data.
15 . The system of claim 14 , wherein the query response data comprises character rendering data based on the media context, and wherein the control circuitry is configured to generate the response media content based on the character rendering data.
16 . The system of claim 14 , wherein the query response data comprises scenery rendering data for scenery based on the media context, and wherein the control circuitry is configured to generate the response media content based on the scenery rendering data.
17 . The system of claim 12 , wherein the query response data comprises response media content generated by the selected response content provider.
18 . The system of claim 17 , wherein the response media content comprises a visual representation of the determined character based on the media context and the query response data.
19 . The system of claim 17 , wherein the response media content comprises visual scenery based on the media context and the query response data.
20 . The system of claim 12 , wherein the control circuitry, when processing the voice query to identify the keyword from the voice query, is configured to:
transmit the textual query to a search engine via the network interface;
receive a search result from the search engine via the network interface; and
identify the keyword from the search result.