Aided system of photography composition
A media application receives, from a server, an identification of a first composition type from a set of compositions to apply to an initial image captured with a user device. Responsive to one or more people being detected in the initial image, the media application generates a modified image, where the one or more people are removed from the initial image to obtain the modified image. The media application scores at least one candidate position within the modified image based on corresponding composition rules for the first composition type. The media application provides a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates a recommended position for the one or more people in the final image.
1 . A computer-implemented method comprising:
providing a geographic location of a user device to a server;
receiving, from the server, a panoramic image that corresponds to the geographic location of the user device, an identification of a first composition type from a set of compositions to apply to an initial image captured with the user device, and image data that describes one or more windows from the panoramic image that are based on the first composition type, wherein the first composition type is selected based on the geographic location of the user device;
responsive to one or more people being detected in the initial image, generating a modified image, wherein the one or more people are removed from the initial image to obtain the modified image;
determining one or more candidate positions within the modified image based on the one or more windows from the panoramic image;
scoring the one or more candidate positions within the modified image based on corresponding composition rules for the first composition type; and
providing a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates one or more recommended positions for the one or more people in the final image based on one or more corresponding scores.
2 . The method of claim 1 , wherein the modified image is generated using pixels from the panoramic image that correspond to locations of pixels in the initial image where the one or more people are removed.
3 . The method of claim 1 , wherein the image data further includes a saliency map of the panoramic image and one or more composition scores that correspond to the one or more windows from the panoramic image.
4 . The method of claim 1 , further comprising:
generating one or more resized versions of the one or more people removed from the initial image by estimating one or more heights of the one or more people in the initial image and one or more distances between the one or more people and a scene being captured by the user device; and
wherein determining the one or more candidate positions is further based on a relative angle of the user device to the scene being captured by the user device and the one or more resized versions of the one or more people.
5 . The method of claim 1 , wherein determining the one or more candidate positions within the modified image is further based on one or more landmarks at the geographic location of the user device.
6 . The method of claim 1 , further comprising:
generating a cropped image from the final image based on the first composition type; and
responsive to the cropped image excluding one or more saliency points, storing the one or more saliency points as metadata associated with the cropped image.
7 . The method of claim 1 , wherein generating the modified image includes:
generating a mask of the one or more people;
removing the mask from the initial image; and
filling in, with pixels, empty space of the initial image that corresponds to the mask.
8 . The method of claim 1 , further comprising adjusting one or more sizes of the one or more people for each candidate position based on one or more distances from the one or more people to a scene being captured in the modified image.
9 . The method of claim 1 , wherein the graphical guide updates as the user moves the user device to include one or more updated recommended positions based on an updated initial image.
10 . A user device comprising:
one or more processors; and
a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:
providing a geographic location of the user device to a server;
receiving, from the server, a panoramic image that corresponds to the geographic location of the user device, coordinates for a first window that is based on a first composition type from a set of compositions to apply to an initial image captured with a camera of the user device, and coordinates for a second window that is based on a second composition type from the set of compositions to apply to the initial image, wherein the first composition type and the second composition type are selected based on the geographic location of the user device;
responsive to one or more people being detected in the initial image generating a modified image, wherein the one or more people are removed from the initial image to obtain the modified image;
determining a first candidate position within the modified image based on the coordinates for the first window and a second candidate position within the modified image based on the coordinates for the second window;
scoring the first candidate position within the modified image based on corresponding composition rules for the first composition type;
scoring the second candidate position within the modified image based on corresponding composition rules for the second composition type; and
providing a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates one or more recommended positions for the one or more people in the final image based on one or more corresponding scores.
11 . The user device of claim 10 , wherein the modified image is generated using pixels from the panoramic image that correspond to locations of pixels in the initial image where the one or more people are removed.
12 . The user device of claim 10 , wherein receiving the panoramic image from the server further includes receiving a saliency map of the panoramic image, a first composition score for the first window from the panoramic image, and a second composition score for the second window from the panoramic image.
13 . The user device of claim 10 , wherein the operations further comprise:
generating one or more resized versions of the one or more people removed from the initial image based on one or more heights of the one or more people and one or more distances between the one or more people and a scene being captured by the user device; and
wherein determining the first candidate position and the second candidate position is further based on a relative angle of the user device to the scene being captured by the user device and the one or more resized versions of the one or more people.
14 . The user device of claim 10 , wherein
determining the first candidate position within the modified image is further based on one or more landmarks at the geographic location of the user device.
15 . The user device of claim 10 , wherein the operations further comprise:
generating a cropped image from the final image based on the first composition type; and
responsive to the cropped image excluding one or more saliency points, storing the one or more saliency points as metadata associated with the cropped image.
16 . A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:
providing a geographic location of a user device to a server;
receiving, from the server, a panoramic image that corresponds to the geographic location of the user device, an identification of a first composition type from a set of compositions to apply to an initial image captured with the user device, and image data that describes one or more windows from the panoramic image that are based on the first composition type, wherein the first composition type is selected based on the geographic location of the user device;
responsive to one or more people being detected in the initial image, generating a modified image, wherein the one or more people are removed from the initial image to obtain the modified image;
determining one or more candidate positions within the modified image based on the one or more windows from the panoramic image;
scoring the one or more candidate positions within the modified image based on corresponding composition rules for the first composition type; and
providing a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates one or more recommended positions for the one or more people in the final image based on one or more corresponding scores.
17 . The non-transitory computer-readable medium of claim 16 , wherein the modified image is generated using pixels from the panoramic image that correspond to locations of pixels in the initial image where the one or more people are removed.
18 . The non-transitory computer-readable medium of claim 16 , wherein the image data further includes a saliency map of the panoramic image and one or more composition scores that correspond to the one or more windows from the panoramic image.
19 . The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:
generating one or more resized versions of the one or more people removed from the initial image by estimating one or more heights of the one or more people in the initial image and one or more distances between the one or more people and a scene being captured by the user device; and
wherein determining the one or more candidate positions is further based on a relative angle of the user device to the scene being captured by the user device and the one or more resized versions of the one or more people.
20 . The non-transitory computer-readable medium of claim 16 , wherein
determining the one or more candidate positions within the modified image is further based on one or more landmarks at the geographic location of the user device.