IP Library › Granted Patent US 12,737,945
Granted Patent B2
US 12,737,945 · App. 18/787,384 · Granted Sep 15, 2026

Generating a group photo that includes a photographer

Inventors: Adi Zicher (Mountain View, CA); Assaf Zomet (Mountain View, CA); Maayan Rossmann (Mountain View, CA); Jung-Chen Hung (Mountain View, CA); Or Guz (Mountain View, CA); Jay Tenenbaum (Mountain View, CA); Avram Golbert (Mountain View, CA); Yaron Brodsky (Mountain View, CA); Fuhao Shi (Mountain View, CA)
Assignee: Google LLC
G06T11/60G06T5/50G06T5/60G06T5/77G06T7/11G06T7/194G06T7/50H04N23/635H04N23/64G06T2200/24G06T2207/20081G06T2207/20212G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,945
App. No.
18/787,384
Granted
Sep 15, 2026
Kind
B2
Abstract

A user device receives a request to generate a composite image. The media application a first image that includes one or more first subjects. The media application determines a previous pose of the user device associated with capture of the first image. The media application segments the one or more first subjects from the first image. The media application generates one or more overlays that correspond to the one or more first subjects based on segmenting the one or more first subjects. The media application displays the one or more overlays on a viewfinder of the user device to provide guidance for a user to capture a second image based on a comparison of a current pose of the user device to the previous pose of the user device. The media application generates the composite image.

Claims (77)

1 . A computer-implemented method comprising:

receiving, at a user device, a request to generate a composite image;

receiving a first image that includes one or more first subjects;

determining a previous pose of the user device associated with capture of the first image;

segmenting the one or more first subjects from the first image;

determining one or more first depths of the one or more first subjects;

generating one or more overlays that correspond to the one or more first subjects based on segmenting the one or more first subjects;

displaying the one or more overlays on a viewfinder of the user device to provide guidance for a user to capture a second image based on a comparison of a current pose of the user device to the previous pose of the user device, wherein the second image includes one or more second subjects; and

generating the composite image that includes the one or more first subjects and the one or more second subjects, wherein the composite image is generated based on the one or more first depths.

2 . The method of claim 1 , further comprising:

determining one or more second depths of the one or more second subjects in the second image; and

determining an order of the one or more first subjects and the one or more second subjects in the composite image based on one or more selected from a group of the one or more first depths, the one or more second depths, and an output from a machine-learning model, wherein the composite image includes the one or more first subjects in front of the one or more second subjects based on the order.

3 . The method of claim 1 , further comprising:

displaying a frame that changes responsive to the comparison of the current pose of the user device to the previous pose of the user device;

wherein the frame is placed at a third depth based on the one or more first depths of the one or more first subjects in the first image; and

wherein the frame includes a width and a height that are in correspondence with the viewfinder.

4 . The method of claim 1 , further comprising:

before generating the one or more overlays, segmenting a background and a foreground of the first image; and

determining that the one or more first subjects are in the foreground, wherein the one or more overlays and the composite image are generated based on the one or more first subjects being in the foreground.

5 . The method of claim 1 , wherein the composite image is a first composite image that includes the one or more segmented first subjects added to the second image and the method further comprises:

determining a first stitching score for the first composite image;

segmenting the one or more second subjects from the second image;

generating a second composite image by adding the one or more segmented second subjects to the first image;

determining a second stitching score for the second composite image; and

providing the first composite image to the user based on the first stitching score being greater than the second stitching score.

6 . The method of claim 1 , wherein segmenting the one or more first subjects from the first image further includes segmenting one or more objects attached to the one or more first subjects.

7 . The method of claim 1 , wherein generating the composite image includes:

generating an intermediate image that combines the one or more segmented first subjects with the second image;

identifying one or more objects that occlude the one or more first subjects or the one or more second subjects in the intermediate image;

responsive to identifying the one or more objects that occlude, determining if one or more gaps are present in the intermediate image; and

responsive to determining that one or more gaps are present, generating the composite image by inpainting the one or more gaps.

8 . The method of claim 1 , wherein generating the composite image includes:

providing the second image as input to a machine-learning model;

outputting, with the machine-learning model, an intermediate image that extends one or more boundaries of the second image and includes filled in pixels between the one or more boundaries of the second image and one or more boundaries of the intermediate image; and

combining the intermediate image with the one or more segmented first subjects to form the composite image.

9 . The method of claim 1 , further comprising:

guiding the user to capture the second image by tilting the user device as compared to the previous pose of the user device associated with capture of the first image.

10 . The method of claim 1 , wherein prior to receiving the request to generate the composite image, the method further comprises:

searching an image library associated with the user device to identify the second image with the one or more second subjects that are missing from the first image; and

responsive to identifying the second image with the one or more second subjects that are missing from the first image, generating a user interface that includes an option to request the composite image to be generated by combining the first image and the second image.

11 . A system comprising:

one or more processors; and

a memory that stores instructions that, when executed by the one or more processors cause the one or more processors to perform operations comprising:

receiving, at a user device, a request to generate a composite image;

receiving a first image that includes one or more first subjects;

determining a previous pose of the user device associated with capture of the first image;

segmenting the one or more first subjects from the first image;

determining one or more first depths of the one or more first subjects;

generating one or more overlays that correspond to the one or more first subjects based on segmenting the one or more first subjects;

displaying the one or more overlays on a viewfinder of the user device to provide guidance for a user to capture a second image based on a comparison of a current pose of the user device to the previous pose of the user device, wherein the second image includes one or more second subjects; and

generating the composite image that includes the one or more first subjects and the one or more second subjects, wherein the composite image is generated based on the one or more first depths.

12 . The system of claim 11 , wherein the operations further include:

determining one or more second depths of the one or more second subjects in the second image; and

determining an order of the one or more first subjects and the one or more second subjects in the composite image based on one or more selected from a group of the one or more first depths, the one or more second depths, and an output from a machine-learning model, wherein the composite image includes the one or more first subjects in front of the one or more second subjects based on the order.

13 . The system of claim 11 , wherein the operations further include:

displaying a frame that changes responsive to the comparison of the current pose of the user device to the previous pose of the user device;

wherein the frame is placed at a third depth based on the one or more first depths of the one or more first subjects in the first image; and

wherein the frame includes a width and a height that are in correspondence with the viewfinder.

14 . A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving, at a user device, a request to generate a composite image;

receiving a first image that includes one or more first subjects;

determining a previous pose of the user device associated with capture of the first image;

segmenting the one or more first subjects from the first image;

determining one or more first depths of the one or more first subjects;

generating one or more overlays that correspond to the one or more first subjects based on segmenting the one or more first subjects;

displaying the one or more overlays on a viewfinder of the user device to provide guidance for a user to capture a second image based on a comparison of a current pose of the user device to the previous pose of the user device, wherein the second image includes one or more second subjects; and

generating the composite image that includes the one or more first subjects and the one or more second subjects, wherein the composite image is generated based on the one or more first depths.

15 . The non-transitory computer-readable medium of claim 14 , wherein the operations further include:

determining one or more second depths of the one or more second subjects in the second image; and

determining an order of the one or more first subjects and the one or more second subjects in the composite image based on one or more selected from a group of the one or more first depths, the one or more second depths, and an output from a machine-learning model, wherein the composite image includes the one or more first subjects in front of the one or more second subjects based on the order.

16 . The non-transitory computer-readable medium of claim 14 , wherein the operations further include:

displaying a frame that changes responsive to the comparison of the current pose of the user device to the previous pose of the user device;

wherein the frame is placed at a third depth based on the one or more first depths of the one or more first subjects in the first image; and

wherein the frame includes a width and a height that are in correspondence with the viewfinder.

17 . The non-transitory computer-readable medium of claim 14 , wherein the operations further include:

before generating the one or more overlays, segmenting a background and a foreground of the first image; and

determining that the one or more first subjects are in the foreground, wherein the one or more overlays and the composite image are generated based on the one or more first subjects being in the foreground.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2024
From: ZICHER, ADI; ZOMET, ASSAF; ROSSMANN, MAAYAN; HUNG, JUNG-CHEN; GUZ, OR; TENENBAUM, JAY; GOLBERT, AVRAM; BRODSKY, YARON; SHI, FUHAO
To: GOOGLE LLC
Reel/Frame 068228/0750 →
Continuity (1)
Related Publication 20260030806A1 · Jan 29, 2026
References Cited (17)
US 7221395B2 · Kinjo · 2007 [cited by examiner]
US 7423671B2 · Kiso · 2008 [cited by examiner]
US 8259187B2 · Liu · 2012 [cited by examiner]
US 9153004B2 · Hirano · 2015 [cited by examiner]
US 9665764B2 · Baek · 2017 [cited by examiner]
US 10121229B2 · Gray · 2018 [cited by applicant]
US 12192616B2 · Ma · 2025 [cited by examiner]
US 20120120273A1 · Amagai et al. · 2012 [cited by applicant]
US 20140184841A1 · Woo et al. · 2014 [cited by applicant]
US 20150009359A1 · Zaheer et al. · 2015 [cited by applicant]
US 20160057363A1 · Posa · 2016 [cited by applicant]
US 20160148428A1 · Agarwal et al. · 2016 [cited by applicant]
US 20170024919A1 · Murphy-Chutorian et al. · 2017 [cited by applicant]
US 20190355172A1 · Dsouza · 2019 [cited by examiner]
US 20210067676A1 · Wen et al. · 2021 [cited by applicant]
US 20240135572A1 · Singh et al. · 2024 [cited by applicant]
WO 2024041394A1 · 2024 [cited by applicant]