IP Library Granted Patent US 12663914
Granted Patent B2
US 12663914 · App. 18/232,146 · Granted Jun 23, 2026

Device-based image modification of depicted objects

Inventors: Theresa Barton (San Mateo, CA); Yanping Chen (San Jose, CA); Jaewook Chung (Mountain View, CA); Christopher Yale Crutchfield (San Diego, CA); Aymeric Damien (San Francisco, CA); Sergei Kotcur (Sochi, RU); Igor Kudriashov (Saratov, RU); Sergey Tulyakov (Santa Monica, CA); Andrew Wan (Marina del Rey, CA); Emre Yamangil (San Francisco, CA)
G06F3/04845G06F3/0482G06T7/11G06V20/40G06V40/161H04N5/2628H04N5/265G06F3/04817G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/20132G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12663914
App. No.
18/232,146
Granted
Jun 23, 2026
Kind
B2
Abstract

A system of machine learning schemes can be configured to efficiently perform image processing tasks on a user device, such as a mobile phone. The system can selectively detect and transform individual regions within each frame of a live streaming video. The system can selectively partition and toggle image effects within the live streaming video.

Claims (57)

1 . A system comprising:

at least one processor;

at least one memory including instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

accessing an image;

processing the image to detect a segment of a first face and a segment of a second face;

segmenting the image to generate the segment of the first face and the segment of the second face;

cropping the image to isolate the segment of the first face and the segment of the second face;

selecting, based on a first facial configuration of the segment of the first face, a first convolutional neural network trained to modify the first facial configuration to have a third facial configuration;

selecting, based on a second facial configuration of the segment of the second face, a second convolutional neural network trained to modify the second facial configuration to have the third facial configuration, the first facial configuration being different than the second facial configuration, and the first convolutional neural network being different than the second convolutional neural network;

processing, by applying the first convolutional neural network on the first face and the second convolutional neural network on the second face, to generate a modified image having a modified segment of the first face and a modified segment of the second face; and

integrating the modified segment of the first face and the modified segment of the second face with the cropped image to generate the modified image.

2 . The system of claim 1 , wherein the first convolutional neural network is a first style transfer neural network trained to transfer images from the first facial configuration to the third facial configuration and the second convolutional neural network is a second style transfer neural network trained to transfer images from the second facial configuration to the third facial configuration.

3 . The system of claim 1 , wherein the operations further comprise:

publishing, to a network site, the modified image as an ephemeral message.

4 . The system of claim 1 , wherein the operations further comprise:

in response to detecting the segment of the first face and the segment of the second face, causing presentation of facial borders, the facial borders outlining the segment of the first face and the segment of the second face within the image; and

causing a user interface to be displayed on a display of the system, the user interface presenting a plurality of third facial configurations.

5 . The system of claim 4 , wherein the plurality of third facial configurations comprises the third facial configuration, and wherein the selecting, based on the first facial configuration of the segment of the first face in the image, is further based on a selection of the third facial configuration by a user.

6 . The system of claim 1 , wherein the image is from a plurality of images, and wherein the plurality of images are a video.

7 . The system of claim 1 , further comprising:

detecting within the image the segment of a third face having a fourth facial configuration; and

refraining from modifying the segment of the third face based on a user not selecting the segment of the third face to be modified.

8 . The system of claim 1 , wherein the first facial configuration, the second facial configuration and the third facial configuration are each at least one of: a smiling facial configuration, a frowning facial configuration, an elder facial configuration, a young facial configuration, or a neutral facial configuration.

9 . The system of claim 1 , wherein the operations further comprise:

normalizing the segment of the first face and the segment of the second face in accordance with a size, a shape, a color, or histogram distribution.

10 . The system of claim 1 , further comprising:

causing the image to be displayed on a display of the system.

11 . A non-transitory computer-readable medium comprising instructions, which when executed by one or more processors, cause the one or more processors to perform operations comprising:

accessing an image;

processing the image to detect a segment of a first face and a segment of a second face;

segmenting the image to generate the segment of the first face and the segment of the second face;

cropping the image to isolate the segment of the first face and the segment of the second face;

selecting, based on a first facial configuration of the segment of the first face, a first convolutional neural network trained to modify the first facial configuration to have a third facial configuration;

selecting, based on a second facial configuration of the segment of the second face, a second convolutional neural network trained to modify the second facial configuration to have the third facial configuration, the first facial configuration being different than the second facial configuration, and the first convolutional neural network being different than the second convolutional neural network;

processing, by applying the first convolutional neural network on the first face and the second convolutional neural network on the second face, to generate a modified image having a modified segment of the first face and a modified segment of the second face; and

integrating the modified segment of the first face and the modified segment of the second face with the cropped image to generate the modified image.

12 . The non-transitory computer-readable medium of claim 11 , wherein the first convolutional neural network is a first style transfer neural network trained to transfer images from the first facial configuration to the third facial configuration and the second convolutional neural network is a second style transfer neural network trained to transfer images from the second facial configuration to the third facial configuration.

13 . The non-transitory computer-readable medium of claim 11 , wherein the image is from a plurality of images, and wherein the plurality of images are a video.

14 . A method comprising:

accessing an image;

processing the image to detect a segment of a first face and a segment of a second face;

segmenting the image to generate the segment of the first face and the segment of the second face;

cropping the image to isolate the segment of the first face and the segment of the second face;

selecting, based on a first facial configuration of the segment of the first face, a first convolutional neural network trained to modify the first facial configuration to have a third facial configuration;

selecting, based on a second facial configuration of the segment of the second face, a second convolutional neural network trained to modify the second facial configuration to have the third facial configuration, the first facial configuration being different than the second facial configuration, and the first convolutional neural network being different than the second convolutional neural network;

processing, by applying the first convolutional neural network on the first face and the second convolutional neural network on the second face, to generate a modified image having a modified segment of the first face and a modified segment of the second face; and

integrating the modified segment of the first face and the modified segment of the second face with the cropped image to generate the modified image.

15 . The method of claim 14 , wherein the first convolutional neural network is a first style transfer neural network trained to transfer images from the first facial configuration to the third facial configuration and the second convolutional neural network is a second style transfer neural network trained to transfer images from the second facial configuration to the third facial configuration.

16 . The method of claim 14 , wherein the method further comprises:

in response to detecting the segment of the first face and the segment of the second face, causing presentation of facial borders, the facial borders outlining the segment of the first face and the segment of the second face within the image; and

causing a user interface to be displayed on a display, the user interface presenting a plurality of third facial configurations.

17 . The method of claim 16 , wherein the plurality of third facial configurations comprises the third facial configuration, and wherein the selecting, based on the first facial configuration of the segment of the first face in the image, is further based on a selection of the third facial configuration by a user.

18 . The method of claim 16 , wherein the image is from a plurality of images, and wherein the plurality of images are a video.

19 . The non-transitory computer-readable medium of claim 11 , wherein the operations further comprise:

in response to detecting the segment of the first face and the segment of the second face, causing presentation of facial borders, the facial borders outlining the segment of the first face and the segment of the second face within the image; and

causing a user interface to be displayed on a display, the user interface presenting a plurality of third facial configurations.

20 . The non-transitory computer-readable medium of claim 19 , wherein the plurality of third facial configurations comprises the third facial configuration, and wherein the selecting, based on the first facial configuration of the segment of the first face in the image, is further based on a selection of the third facial configuration by a user.