Generating image content
Systems and techniques are described herein for generating image content. For instance, an apparatus for generating image content is provided. The method may include a user interface configured to: display an image of a field of view; and receive a user input indicative of a desired change to the image, wherein the desired change to the image comprises a change to the field of view; and at least one processor configured to: provide at least part of the image and an indication of the desired change as inputs to a generative machine-learning model; obtain an altered image from the generative machine-learning model, wherein the altered image comprises at least part of the image of the field of view and generated pixels outside of the field of view.
1 . An apparatus for generating image content, the apparatus comprising:
a user interface comprising a touch screen, the user interface configured to:
display an image of a field of view; and
receive a touch input indicative of a desired change to the image, wherein the desired change to the image comprises a change to the field of view from the field of view to a changed field of view, wherein the changed field of view comprises a portion outside the field of view; and
at least one processor configured to:
provide at least part of the image and an indication of the desired change as inputs to a generative machine-learning model;
obtain an altered image from the generative machine-learning model, wherein the altered image comprises at least part of the image of the field of view and additional pixel data representing the portion outside the field of view.
2 . The apparatus of claim 1 , wherein the user interface is configured to interpret a gesture to receive the touch input.
3 . The apparatus of claim 2 , wherein the user interface is configured to interpret the touch input as the gesture.
4 . The apparatus of claim 3 , wherein, the user interface is configured to interpret a pinch touch gesture as an indication that the desired change comprises expanding the field of view.
5 . The apparatus of claim 3 , wherein the user interface is configured to interpret a drag touch gesture as an indication that the desired change comprises panning to change the field of view.
6 . The apparatus of claim 3 , wherein the user interface is configured to interpret a rotating touch gesture as an indication that the desired change comprises rotating the field of view.
7 . The apparatus of claim 2 , wherein the user interface comprises a camera configured to capture images of a user and wherein the user interface is configured to interpret a pose of the user as the gesture.
8 . The apparatus of claim 7 , wherein the camera comprises at least one of:
an active-depth camera;
an infrared (IR) camera;
a red-green-blue (RGB) camera;
stereo cameras; or
an eye-facing camera.
9 . The apparatus of claim 1 , wherein the altered image comprises the additional pixel data on at least two sides of the at least part of the image of the field of view.
10 . The apparatus of claim 1 , wherein the altered image comprises the additional pixel data on at least one side of the at least part of the image of the field of view.
11 . The apparatus of claim 1 , wherein the at least one processor is further configured to generate a final image by at least one of rotating or cropping the altered image based on the change to the field of view.
12 . The apparatus of claim 1 , wherein the at least one processor is further configured to at least one of rotate or crop the image to obtain the at least part of the image to provide to the generative machine-learning model.
13 . The apparatus of claim 1 , wherein:
the image comprises a first image;
the apparatus further comprises a first camera configured to capture the first image;
the apparatus further comprises a second camera configured to capture a second image; and
the at least one processor is configured to provide at least part of the first image and at least part of the second image as inputs to the generative machine-learning model.
14 . The apparatus of claim 13 , wherein the first camera has a first focal length and the second camera has a second focal length that is different than the first focal length.
15 . The apparatus of claim 1 , wherein:
the image comprises a first image of a scene captured at a first time; and
the at least one processor is further configured to provide at least part of the first image and at least part of a second image of the scene captured at a second time as inputs to the generative machine-learning model.
16 . The apparatus of claim 15 , wherein:
the field of view comprises a first field of view of the scene;
the first image is of the first field of view of the scene; and
the second image is of a second field of view of the scene.
17 . The apparatus of claim 15 , wherein the at least one processor is further configured to determine to use the second image based on the desired change to the first image.
18 . The apparatus of claim 1 , wherein:
prior to the at least one processor obtaining the altered image the at least one processor is configured to obtain additional pixels and cause the user interface to display the image of the field of view and the additional pixels outside the field of view; and
responsive to the at least one processor obtaining the altered image, the at least one processor is configured to cause the user interface to display the altered image.
19 . The apparatus of claim 18 , wherein the additional pixels are blurry.
20 . The apparatus of claim 1 , wherein the user interface is further configured to display the altered image.
21 . The apparatus of claim 1 , wherein the at least one processor is further configured to cause the altered image to be at least one of: displayed by the user interface, stored at a memory, analyzed, or transmitted.
22 . The apparatus of claim 1 , wherein the at least one processor is further configured to smooth pixels at edges between the at least part of the image of the field of view and the additional pixel data outside the field of view.
23 . The apparatus of claim 1 , wherein the at least one processor implements the generative machine-learning model.
24 . The apparatus of claim 1 , further comprising a communication interface configured to:
transmit at least part of the image and the indication of the desired change to a computing device that implements the generative machine-learning model; and
receive the altered image from the computing device.
25 . The apparatus of claim 1 , wherein the apparatus further comprises an orientation sensor configured to sense an orientation of the apparatus and wherein the user interface is configured to interpret a rotation of the apparatus as an indication that the desired change comprises rotating the field of view.
26 . The apparatus of claim 1 , wherein one of:
the image is displayed by the user interface in a portrait format and the altered image comprises the additional pixel data on at least a side of the at least part of the image of the field of view; or
the image is displayed by the user interface in a landscape format and the altered image comprises the additional pixel data at least one of above or below the at least part of the image of the field of view.
27 . The apparatus of claim 1 , wherein the altered image comprises the additional pixel data at at least one corner of the altered image.
28 . A method comprising for generating image content, the method comprising:
displaying an image of a field of view at a touch screen of a user interface;
receiving, at the user interface, a touch input indicative of a desired change to the image, wherein the desired change to the image comprises a change to the field of view from the field of view to a changed field of view, wherein the changed field of view comprises a portion outside the field of view;
providing at least part of the image and an indication of the desired change as inputs to a generative machine-learning model; and
obtaining an altered image from the generative machine-learning model, wherein the altered image comprises at least part of the image of the field of view and additional pixel data representing the portion outside the field of view.