User directed video generation method and system
A user directed video generation method and system obtains a natural language-based communication from a user requesting that a computer-implemented system generate a virtual environment that is based on a description that is provided by the user. The description is interpreted by a trained neural network. Representations of pixel patterns are generated by a trained neural network in accordance with the interpretation. The representations of the pixel patterns are evaluated for consistency with context and then selected based on the evaluation. The selected pixel patterns are embodied in a video stream that is provided to the user. Natural language that may be in audio form may be generated to accompany the video stream.
1 . A computer-implemented method, comprising:
obtaining a natural language-based communication from a user requesting that a computer-implemented system generate an alternative reality that is based on a description provided by the user, wherein the description comprises natural language that comprises a plurality of syntactical elements;
interpreting automatically the description by applying a first trained computer-implemented neural network to interpret the plurality of syntactical elements;
generating automatically by applying a second trained computer-implemented neural network a plurality of pixel patterns, wherein each of the plurality of pixel patterns are generated based on the interpreting of the description;
evaluating automatically the each of the plurality of the pixel patterns, wherein the evaluating is with respect to a context;
selecting automatically, based on the evaluating of the each of the plurality of the pixel patterns, one or more of the pixel patterns from the each of the plurality of the pixel patterns;
generating automatically a video stream that embodies the description, wherein the video stream includes one or more pixel patterns that each correspond to the selected one or more of the pixel patterns; and
providing the video stream to the user.
2 . The method of claim 1 , wherein at least one of the plurality of the pixel patterns has a trained correspondence with one or more syntactical elements.
3 . The method of claim 1 , wherein at least one of the plurality of the pixel patterns represents an object in motion.
4 . The method of claim 1 , wherein the evaluating with respect to the context comprises determining a probability that the each of the plurality of the pixel patterns is consistent with the context.
5 . The method of claim 1 , wherein the context comprises a physical environment.
6 . The method of claim 1 , wherein the each of the plurality of the pixel patterns has an associated spatial and temporal indicator.
7 . The method of claim 1 , wherein the second trained computer-implemented neural network operates with an adversarial third trained computer-implemented neural network.
8 . The method of claim 1 , wherein the video stream is further generated in accordance with an inference of a preference of the user that is based on a plurality of usage behaviors that occur prior to the user requesting that the computer-implemented system generate the virtual environment.
9 . A computer-implemented system comprising one or more processor-based devices configured to:
obtain a natural language-based communication from a user requesting that the computer-implemented system generate an alternative reality that is based on a description provided by the user, wherein the description comprises natural language that comprises a plurality of syntactical elements;
interpret automatically the description by applying a first trained computer-implemented neural network to interpret the plurality of syntactical elements;
generate automatically by applying a second trained computer-implemented neural network a plurality of pixel patterns, wherein each of the plurality of pixel patterns are generated based on the interpreting of the description;
evaluate automatically the each of the plurality of the pixel patterns, wherein the evaluation is with respect to a context;
select automatically, based on the evaluating of the each of the plurality of the pixel patterns, one or more of the pixel patterns from the each of the plurality of the plurality of the pixel patterns;
generate automatically a video stream that embodies the description, wherein the video stream includes one or more pixel patterns that each correspond to the selected one or more of the pixel patterns; and
provide the video stream to the user.
10 . The computer-implemented system of claim 9 , wherein at least one of the plurality of the pixel patterns has a trained correspondence with one or more syntactical elements.
11 . The computer-implemented system of claim 9 , wherein at least one of the plurality of the pixel patterns represents an object in motion.
12 . The computer-implemented system of claim 9 , wherein the evaluation with respect to the context comprises determining a probability that the each of the plurality of the pixel patterns is consistent with the context.
13 . The computer-implemented system of claim 9 , wherein the evaluation of the each of the plurality of the pixel patterns comprises further evaluating if the each of the plurality of the pixel patterns is in accordance with a scenario.
14 . The computer-implemented system of claim 9 , wherein the each of the plurality of the pixel patterns has an associated spatial and temporal indicator.
15 . The computer-implemented system of claim 9 , wherein the second trained computer-implemented neural network operates with an adversarial third trained computer-implemented neural network.
16 . The computer-implemented system of claim 9 , wherein the video stream is further generated in accordance with an inference of a preference of the user that is based on a plurality of usage behaviors that occur prior to the user requesting that the computer-implemented system generate the virtual environment.
17 . A computer-implemented system comprising one or more processor-based devices configured to:
obtain a natural language-based communication from a user requesting that the computer-implemented system generate an alternative reality that is based on a description provided by the user, wherein the description comprises natural language that comprises a plurality of syntactical elements;
interpret automatically the description by applying a first trained computer-implemented neural network to interpret the plurality of syntactical elements;
generate automatically by applying a second trained computer-implemented neural network a plurality of pixel patterns, wherein each of the plurality of pixel patterns correspond to one of a plurality of pixel patterns and are generated based on the interpreting of the description;
evaluate automatically the each of the plurality of the pixel patterns, wherein the evaluating is with respect to a context;
select automatically, based on the evaluating of the plurality of representations of the pixel patterns, one or more pixel patterns from the plurality of the pixel patterns;
generate automatically a video stream and an associated plurality of syntactical elements that embody the description, wherein the video stream includes one or more pixel patterns that each correspond to the selected one or more of the pixel patterns and the included one or more pixel patterns each have corresponding one or more syntactical elements of the associated plurality of syntactical elements; and
provide the video stream and the associated plurality of syntactical elements to the user.
18 . The computer-implemented system of claim 17 , wherein the associated plurality of syntactical elements comprises natural language that is generated by a trained computer-implemented neural network.
19 . The computer-implemented system of claim 18 , wherein the associated plurality of syntactical elements is provided to the user in audio form.
20 . The computer-implemented system of claim 18 , wherein each of the corresponding one or more syntactical elements of the associated plurality of syntactical elements is generated based on probabilistic-based correspondences with the included one or more pixel patterns.