Multi-modal user inputs for phrase building
Techniques for generating and providing phrases built using multi-modal inputs are described herein. These techniques may include generating one or more phrase trees for use in parsing options for user selections through multi-modal inputs to iteratively build an input phrase such as an open-ended phrase. The multi-modal inputs may include gaze inputs, touch inputs, and other inputs to select terms for building a phrase. The selected terms may be used to build a phrase using trees of terms arranged in sequential nodes based on frequently used input phrases to a model.
1 . A system comprising:
a display;
a camera directed towards a user space configured to capture image data of a user interacting with the display;
one or more processors; and
one or more non-transitory computer-readable media storing computer-executable instructions that, when executed on the one or more processors, cause the one or more processors to perform acts comprising:
generating, for presentation on the display, a first user interface;
determining a user selection of a graphical display element of the first user interface associated with providing a text input to the system;
determining, based on user data, a first set of phrase trees to download from a cloud-based storage system, the first set of phrase trees comprising terms and sets of possible subsequent terms for building one or more phrases, the first set of phrase trees generated based on instructions provided to a virtual assistant;
downloading, to the system, limited portions of phrase trees of the first set of phrase trees, the limited portions of the phrase trees comprising limited layers of nodes of respective ones of the phrase trees;
determining, based on the user data and using the limited portions of the phrase trees downloaded to the system, a first set of terms for presentation within a second user interface on the display, the first set of terms comprising initial terms of one or more of the phrase trees;
receiving, from the camera, first image data;
determining, based on the first image data, a first gaze location within the second user interface;
determining, based on the first gaze location, a first term of the first set of terms;
determining, based on the first term and the limited portions of the phrase trees downloaded to the system, one or more subsequent terms for presentation within a third user interface on the display;
receiving, from the camera, second image data;
determining, based on the second image data, a second gaze location on the third user interface;
determining, based on the second gaze location and the third user interface, a second term; and
determining a phrase for input into the system based on the first term and the second term.
2 . The system of claim 1 , wherein the acts further comprise:
determining, based on the user data and the phrase including the first term and the second term, a second set of phrase trees to download from the cloud-based storage system;
determining, based on the first term and the second term, one or more subsequent terms to display on a fourth user interface;
receiving, from the camera, third image data;
determining, based on the third image data, a third gaze location on the fourth user interface; and
determining, based on the third gaze location and the fourth user interface, a third term, wherein determining the phrase is further based on the third term.
3 . The system of claim 1 , wherein the phrase comprises a phrase for interacting with a digital voice assistant.
4 . The system of claim 1 , wherein the display comprises at least one of:
a display of a tablet device;
a display of an augmented reality device;
a display of a virtual reality device; or
a display of a computing device.
5 . A method comprising:
generating, for presentation on a display of a user device, a first set of user interface elements;
determining a user selection of a first element of the first set of user interface elements;
determining, in response to the user selection of the first element, a first set of phrase trees to download from a cloud-based storage system, the first set of phrase trees comprising initial terms and sets of possible subsequent terms for building one or more phrases, and wherein the first set of phrase trees are derived based on a set of potential inputs to a voice-activated system;
downloading, to the user device, limited portions of phrase trees of the first set of phrase trees, the limited portions of the phrase trees comprising limited layers of nodes of respective ones of the phrase trees;
determining, based on the first element and using the limited portions of the phrase trees downloaded to the user device, a first set of terms to display with a second set of user interface elements, the first set of terms comprising initial terms of one or more of the phrase trees;
determining, based on a first user input, a first term of the first set of terms, wherein the first user input comprises a non-speech input;
determining, based on the first term and the limited portions of the phrase trees downloaded to the user device, one or more subsequent terms to display with a third set of user interface elements;
determining, based on a second user input and the third set of user interface elements, a second term of the one or more subsequent terms; and
determining a phrase based on the first term and the second term.
6 . The method of claim 5 , further comprising:
determining, based on the first term and the second term, a second set of phrase trees to download from the cloud-based storage system, the second set of phrase trees comprising second terms and second sets of possible subsequent terms;
determining, based on the second set of phrase trees, the first term, and the second term, a third set of terms to display with a fourth set of user interface elements; and
determining, based on a third user input, a third term of the third set of terms, wherein determining the phrase is further based on the third term.
7 . The method of claim 5 , wherein the first set of phrase trees comprise phase trees including up to five levels of subsequent terms.
8 . The method of claim 5 , wherein the first set of phrase trees are based at least in part on a database of frequent user input phrases to the voice-activated system.
9 . The method of claim 8 , wherein the database of frequent user input phrases comprises user input phrases across a plurality of users.
10 . The method of claim 8 , wherein the database of frequent user input phrases comprises user input phrases based on previous inputs of the user.
11 . The method of claim 5 , wherein the first user input comprises a first gaze input and the second user input comprises a second gaze input.
12 . The method of claim 5 , wherein the first user input and the second user input comprise selection of a graphical element on a display of a user device.
13 . The method of claim 5 , wherein determining the first set of terms to display and the one or more subsequent terms comprises providing receiving an output of a predictive text model.
14 . One or more non-transitory computer-readable media storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform acts comprising:
generating, for presentation on a display of a user device, a first set of user interface elements;
determining a user selection of a first element of the first set of user interface elements;
determining, in response to the user selection of the first element, a first set of phrase trees to download from a cloud-based storage system, the first set of phrase trees comprising initial terms and sets of possible subsequent terms for building one or more phrases, and wherein the first set of phrase trees are derived based on a set of potential inputs to a voice-activated system;
downloading, to the user device, limited portions of phrase trees of the first set of phrase trees, the limited portions of the phrase trees comprising limited layers of nodes of respective ones of the phrase trees;
determining, based on the first element and using the limited portions of the phrase trees downloaded to the user device, a first set of terms to display with a second set of user interface elements, the first set of terms comprising initial terms of one or more of the phrase trees;
determining, based on a first user input, a first term of the first set of terms, wherein the first user input comprises a non-speech input;
determining, based on the first term and the limited portions of the phrase trees downloaded to the user device, one or more subsequent terms to display with a third set of user interface elements;
determining, based on a second user input and the third set of user interface elements, a second term of the one or more subsequent terms; and
determining a phrase based on the first term and the second term.
15 . The one or more non-transitory computer-readable media of claim 14 , wherein the acts further comprise:
determining, based on the first term and the second term, a second set of phrase trees to download from the cloud-based storage system, the second set of phrase trees comprising second terms and second sets of possible subsequent terms;
determining, based on the second set of phrase trees, the first term, and the second term, a third set of terms to display with a fourth set of user interface elements; and
determining, based on a third user input, a third term of the third set of terms, wherein determining the phrase is further based on the third term.
16 . The one or more non-transitory computer-readable media of claim 14 , wherein determining the one or more subsequent terms to display with the third set of user interface elements comprises providing the first term and the phrase tree to a predictive model.
17 . The one or more non-transitory computer-readable media of claim 14 , wherein the first set of phrase trees are based at least in part on a database of frequent user input phrases to the voice-activated system.
18 . The one or more non-transitory computer-readable media of claim 14 , wherein determining the first set of phase trees is further based on an environment surrounding a user device.
19 . The one or more non-transitory computer-readable media of claim 14 , wherein the acts further comprise processing the phrase to perform one or more acts via a cloud-based computing environment.
20 . The one or more non-transitory computer-readable media of claim 14 , wherein the first user input comprises a gaze input determined by using a machine learning algorithm trained to determine gaze location on a display based on image data from a camera associated with the display.