Interactive content output
Techniques for outputting interactive content and processing interactions with respect to the interactive content are described. While outputting requested content, a system may determine that interactive content is to be outputted. The system may determine output data including a first portion indicating that interactive content is going to be output and a second portion representing content corresponding to an item. The system may send the output data to the device. A user may interact with the output data, for example, by requesting performance of an action with respect to the item.
1 . A computer-implemented method, comprising:
causing a device to output first content;
determining second content is to be output, the second content corresponding to a first item;
generating first output data comprising:
a first portion representing interactive content is going to be output, and
a second portion corresponding to the second content;
sending the first output data to the device;
receiving, from the device, input data representing a natural language input;
performing natural language processing on the input data to determine natural language processing results data;
determining the natural language processing results data corresponds to the interactive content; and
performing an action responsive to the natural language input.
2 . The computer-implemented method of claim 1 , wherein the second content corresponds to pre-recorded content.
3 . The computer-implemented method of claim 1 , further comprising:
determining a duration of time elapsed between output of the first output data at the device and receipt of the input data;
based on the duration of time, determining that the natural language input corresponds to the first output data; and
sending data corresponding to the input data to an interactive content component for processing.
4 . The computer-implemented method of claim 1 , wherein:
the input data comprises first audio data representing a spoken input; and
performing natural language processing on the input data comprises:
determining automatic speech recognition (ASR) data corresponding to the first audio data, and
determining, using the ASR data, natural language understanding (NLU) data corresponding to the first audio data.
5 . The computer-implemented method of claim 1 , further comprising:
determining context data indicating that the first output data was outputted at the device; and
based at least in part on the context data, determining that the natural language input corresponds to the second content.
6 . The computer-implemented method of claim 1 , further comprising:
processing the natural language processing results data using an interactive content component to determine the action.
7 . The computer-implemented method of claim 1 , further comprising:
determining, using stored data associated with the second content, that a first application component is to be invoked to perform the action; and
sending, to the first application component, a first indication of the action and a second indication of the first item.
8 . The computer-implemented method of claim 1 , further comprising:
prior to sending the first output data, receiving, from the device, second input data requesting output of the first content;
determining second output data corresponding to the first content, the second output data including a time condition when interactive content can be output;
sending, to the device, the second output data;
after a portion of the second output data is outputted, determining that the time condition is satisfied; and
sending, to the device, the first output data in response to the time condition being satisfied.
9 . The computer-implemented method of claim 1 , further comprising:
determining an account identifier corresponding to the device; and
storing data associating the action and the second content to the account identifier.
10 . The computer-implemented method of claim 1 , further comprising:
generating the first output data to comprise a third portion representing that the action can be performed with respect to the first item.
11 . A system comprising:
at least one processor; and
at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
cause a device to output first content;
determine second content is to be output, the second content corresponding to a first item;
generate first output data comprising:
a first portion representing interactive content is going to be output, and
a second portion corresponding to the second content;
send the first output data to the device;
receive, from the device, input data representing a natural language input;
perform natural language processing on the input data to determine natural language processing results data;
determine the natural language processing results data corresponds to the interactive content; and
perform an action responsive to the natural language input.
12 . The system of claim 11 , wherein the second content corresponds to pre-recorded content.
13 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine a duration of time elapsed between output of the first output data at the device and receipt of the input data;
based on the duration of time, determine that the natural language input corresponds to the first output data; and
send data corresponding to the input data to an interactive content component for processing.
14 . The system of claim 11 , wherein:
the input data comprises first audio data representing a spoken input; and
the instructions that cause the system to natural language processing on the input data comprise instructions that, when executed by the at least one processor, cause the system to:
determine automatic speech recognition (ASR) data corresponding to the first audio data, and
determine, using the ASR data, natural language understanding (NLU) data corresponding to the first audio data.
15 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine context data indicating that the first output data was outputted at the device; and
based at least in part on the context data, determine that the natural language input corresponds to the second content.
16 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the natural language processing results data using an interactive content component to determine the action.
17 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine, using stored data associated with the second content, that a first application component is to be invoked to perform the action; and
send, to the first application component, a first indication of the action and a second indication of the first item.
18 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
prior to sending the first output data, receive, from the device, second input data requesting output of the first content;
determine second output data corresponding to the first content, the second output data including a time condition when interactive content can be output;
send, to the device, the second output data;
after a portion of the second output data is outputted, determine that the time condition is satisfied; and
send, to the device, the first output data in response to the time condition being satisfied.
19 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine an account identifier corresponding to the device; and
store data associating the action and the second content to the account identifier.
20 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
generate the first output data to comprise a third portion representing that the action can be performed with respect to the first item.