IP Library › Granted Patent US 12,634,575
Granted Patent B2
US 12,634,575 · App. 18/293,689 · Granted May 19, 2026

Aided system of photography composition

Inventor: John Chang (Mountain View, CA)
Assignee: Google LLC
H04N23/64H04N23/633
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,575
App. No.
18/293,689
Granted
May 19, 2026
Kind
B2
Abstract

A media application receives, from a server, an identification of a first composition type from a set of compositions to apply to an initial image captured with a user device. Responsive to one or more people being detected in the initial image, the media application generates a modified image, where the one or more people are removed from the initial image to obtain the modified image. The media application scores at least one candidate position within the modified image based on corresponding composition rules for the first composition type. The media application provides a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates a recommended position for the one or more people in the final image.

Claims (56)

1 . A computer-implemented method comprising:

providing a geographic location of a user device to a server;

receiving, from the server, a panoramic image that corresponds to the geographic location of the user device, an identification of a first composition type from a set of compositions to apply to an initial image captured with the user device, and image data that describes one or more windows from the panoramic image that are based on the first composition type, wherein the first composition type is selected based on the geographic location of the user device;

responsive to one or more people being detected in the initial image, generating a modified image, wherein the one or more people are removed from the initial image to obtain the modified image;

determining one or more candidate positions within the modified image based on the one or more windows from the panoramic image;

scoring the one or more candidate positions within the modified image based on corresponding composition rules for the first composition type; and

providing a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates one or more recommended positions for the one or more people in the final image based on one or more corresponding scores.

2 . The method of claim 1 , wherein the modified image is generated using pixels from the panoramic image that correspond to locations of pixels in the initial image where the one or more people are removed.

3 . The method of claim 1 , wherein the image data further includes a saliency map of the panoramic image and one or more composition scores that correspond to the one or more windows from the panoramic image.

4 . The method of claim 1 , further comprising:

generating one or more resized versions of the one or more people removed from the initial image by estimating one or more heights of the one or more people in the initial image and one or more distances between the one or more people and a scene being captured by the user device; and

wherein determining the one or more candidate positions is further based on a relative angle of the user device to the scene being captured by the user device and the one or more resized versions of the one or more people.

5 . The method of claim 1 , wherein determining the one or more candidate positions within the modified image is further based on one or more landmarks at the geographic location of the user device.

6 . The method of claim 1 , further comprising:

generating a cropped image from the final image based on the first composition type; and

responsive to the cropped image excluding one or more saliency points, storing the one or more saliency points as metadata associated with the cropped image.

7 . The method of claim 1 , wherein generating the modified image includes:

generating a mask of the one or more people;

removing the mask from the initial image; and

filling in, with pixels, empty space of the initial image that corresponds to the mask.

8 . The method of claim 1 , further comprising adjusting one or more sizes of the one or more people for each candidate position based on one or more distances from the one or more people to a scene being captured in the modified image.

9 . The method of claim 1 , wherein the graphical guide updates as the user moves the user device to include one or more updated recommended positions based on an updated initial image.

10 . A user device comprising:

one or more processors; and

a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:

providing a geographic location of the user device to a server;

receiving, from the server, a panoramic image that corresponds to the geographic location of the user device, coordinates for a first window that is based on a first composition type from a set of compositions to apply to an initial image captured with a camera of the user device, and coordinates for a second window that is based on a second composition type from the set of compositions to apply to the initial image, wherein the first composition type and the second composition type are selected based on the geographic location of the user device;

responsive to one or more people being detected in the initial image generating a modified image, wherein the one or more people are removed from the initial image to obtain the modified image;

determining a first candidate position within the modified image based on the coordinates for the first window and a second candidate position within the modified image based on the coordinates for the second window;

scoring the first candidate position within the modified image based on corresponding composition rules for the first composition type;

scoring the second candidate position within the modified image based on corresponding composition rules for the second composition type; and

providing a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates one or more recommended positions for the one or more people in the final image based on one or more corresponding scores.

11 . The user device of claim 10 , wherein the modified image is generated using pixels from the panoramic image that correspond to locations of pixels in the initial image where the one or more people are removed.

12 . The user device of claim 10 , wherein receiving the panoramic image from the server further includes receiving a saliency map of the panoramic image, a first composition score for the first window from the panoramic image, and a second composition score for the second window from the panoramic image.

13 . The user device of claim 10 , wherein the operations further comprise:

generating one or more resized versions of the one or more people removed from the initial image based on one or more heights of the one or more people and one or more distances between the one or more people and a scene being captured by the user device; and

wherein determining the first candidate position and the second candidate position is further based on a relative angle of the user device to the scene being captured by the user device and the one or more resized versions of the one or more people.

14 . The user device of claim 10 , wherein

determining the first candidate position within the modified image is further based on one or more landmarks at the geographic location of the user device.

15 . The user device of claim 10 , wherein the operations further comprise:

generating a cropped image from the final image based on the first composition type; and

responsive to the cropped image excluding one or more saliency points, storing the one or more saliency points as metadata associated with the cropped image.

16 . A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

providing a geographic location of a user device to a server;

receiving, from the server, a panoramic image that corresponds to the geographic location of the user device, an identification of a first composition type from a set of compositions to apply to an initial image captured with the user device, and image data that describes one or more windows from the panoramic image that are based on the first composition type, wherein the first composition type is selected based on the geographic location of the user device;

responsive to one or more people being detected in the initial image, generating a modified image, wherein the one or more people are removed from the initial image to obtain the modified image;

determining one or more candidate positions within the modified image based on the one or more windows from the panoramic image;

scoring the one or more candidate positions within the modified image based on corresponding composition rules for the first composition type; and

providing a graphical guide on a viewfinder of the user device to guide a user to capture a final image, wherein the graphical guide indicates one or more recommended positions for the one or more people in the final image based on one or more corresponding scores.

17 . The non-transitory computer-readable medium of claim 16 , wherein the modified image is generated using pixels from the panoramic image that correspond to locations of pixels in the initial image where the one or more people are removed.

18 . The non-transitory computer-readable medium of claim 16 , wherein the image data further includes a saliency map of the panoramic image and one or more composition scores that correspond to the one or more windows from the panoramic image.

19 . The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

generating one or more resized versions of the one or more people removed from the initial image by estimating one or more heights of the one or more people in the initial image and one or more distances between the one or more people and a scene being captured by the user device; and

wherein determining the one or more candidate positions is further based on a relative angle of the user device to the scene being captured by the user device and the one or more resized versions of the one or more people.

20 . The non-transitory computer-readable medium of claim 16 , wherein

determining the one or more candidate positions within the modified image is further based on one or more landmarks at the geographic location of the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2024
From: CHANG, JOHN
To: GOOGLE LLC
Reel/Frame 066313/0753 →
Continuity (1)
Related Publication 20240334043A1 · Oct 3, 2024
References Cited (16)
US 9626584B2 · Lin et al. · 2017 [cited by applicant]
US 9667860B2 · Hakim et al. · 2017 [cited by applicant]
US 9727802B2 · Wang et al. · 2017 [cited by applicant]
US 10924661B2 · Vasconcelos et al. · 2021 [cited by applicant]
US 11100325B2 · Xu et al. · 2021 [cited by applicant]
US 20150201130A1 · Cho · 2015 [cited by examiner]
US 20170111574A1 · Miyashita · 2017 [cited by applicant]
US 20170193324A1 · Chen et al. · 2017 [cited by applicant]
US 20210110589A1 · Zhang et al. · 2021 [cited by applicant]
US 20240284041A1 · Kokubo · 2024 [cited by examiner]
CN 109146892 · 2018 [cited by applicant]
CN 111327829 · 2021 [cited by applicant]
EPO, International Search Report for International Patent Application No. PCT/US2022/036033, Mar. 24, 2023, 4 pages. [cited by applicant]
EPO, Written Opinion for International Patent Application No. PCT/US2022/036033, Mar. 24, 2023, 8 pages. [cited by applicant]
Greco, Luca et al., “Saliency based aesthetic cut of digital images”, International conference on image analysis and processing. Springer, Berlin, Heidelberg, 2013, 10 pages. [cited by applicant]
Rawat, Yogesh Singh et al., “Context-Aware Photography Learning for Smart Mobile Devices”, ACM Transactions on Multimedia Computing Communications and Applications, Association for Computing Machinery (US), vol. 12, No.… [cited by applicant]