IP Library Granted Patent US 11,829,854
Granted Patent B2
US 11,829,854 · App. 17/403,804 · Granted Nov 28, 2023

Using machine learning to detect which part of the screen includes embedded frames of an uploaded video

Inventors: Filip Pavetic (Zürich, CH); King Hong Thomas Leung (Saratoga, CA); Dmitrii Tochilkin (Zürich, CH)
Assignee: Google LLC
G06N20/00G06F18/214G06F18/241G06V10/25G06V10/764G06V20/40G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,829,854
App. No.
17/403,804
Granted
Nov 28, 2023
Kind
B2
Abstract

A system and methods are disclosed for using a trained machine learning model to identify constituent images within composite images. A method may include providing pixel data of a first image as input to the trained machine learning model, obtaining one or more outputs from the trained machine learning model, and extracting, from the one or more outputs, an indication that the first image is a composite image that includes a constituent image, wherein at least a portion of the constituent image is in a spatial area of the first image.

Claims (42)

1. A system comprising:

a memory to store instructions; and

a processing device to execute the instructions to perform operations comprising:

generating training data for a machine learning model, wherein the training data comprises:

pixel data of a composite image having a first portion containing pixel data of a first constituent image, and a second portion containing pixel data of a second constituent image, the first portion being located at a particular position within the composite image; and

providing the training data to train the machine learning model on pixel data of a plurality of composite images including the composite image,

wherein the trained machine learning model is to receive a new image as input and to produce a new output based on the new image, the new output indicating whether the new image is a composite image containing a constituent image.

2. The system of claim 1 wherein the second portion of the composite image surrounds the first portion of the composite image.

3. The system of claim 1 wherein the first constituent image is a frame of a first video and the second constituent image is a frame of a second video.

4. The system of claim 1 wherein the position of the first portion within the composite image comprises coordinates of an upper left corner of the first constituent image and coordinates of a lower right corner of the first constituent image.

5. The system of claim 1 wherein the training data includes, for each of the plurality of composite images, pixel data of a corresponding composite image, a position of a first portion within the corresponding composite image, and a mapping of the pixel data of the corresponding composite image to the position of the first portion within the corresponding composite image.

6. The system of claim 1 wherein the new output indicates (i) a level of confidence that the new image is a composite image including a constituent image, and (ii) a spatial area in which the constituent image is located within the new image.

7. A method comprising:

providing pixel data of a first image as input to a machine learning model trained using training data comprising pixel data of a plurality of composite images that each include pixel data of respective constituent images;

obtaining one or more outputs from the trained machine learning model; and

extracting, from the one or more outputs, an indication that the first image is a composite image that includes a constituent image, wherein at least a portion of the constituent image is in a spatial area of the first image.

8. The method of claim 7 , wherein:

the indication that the first image is a composite image including a constituent image comprises a level of confidence value; and

the method further comprises:

determining that the level of confidence value satisfies a threshold condition; and

extracting the constituent image from the spatial area of the first image.

9. The method of claim 7 wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein the plurality of spatial areas are uniform in size and shape.

10. The method of claim 7 wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas have different sizes.

11. The method of claim 7 wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas overlap.

12. The method of claim 7 wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein the plurality of spatial areas are non-overlapping.

13. The method of claim 7 wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas have different shapes.

14. The method of claim 7 wherein the training data comprises an input-output mapping comprising an input and an output, the input based on pixel data of a composite image, the composite image comprising a first portion containing pixel data of a second image and a second portion containing pixel data of a third image, and the output identifying a position of the first portion within the composite image.

15. A non-transitory computer readable medium comprising instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving an input image;

processing the input image using a machine learning model trained based on training data comprising pixel data of a plurality of composite images that each include pixel data of respective constituent images; and

obtaining, based on the processing of the input image using the trained machine learning model, one or more outputs comprising an indication that the input image is a composite image including a constituent image, wherein at least a portion of the constituent image is in a spatial area of the input image.

16. The non-transitory computer readable medium of claim 15 , wherein:

the indication that the input image is a composite image including a constituent image comprises a level of confidence value; and

the operations further comprise:

determining that the level of confidence value satisfies a threshold condition; and

extracting the constituent image from the spatial area within the input image.

17. The non-transitory computer readable medium of claim 15 wherein the input image comprises a second constituent image that surrounds the constituent image.

18. The non-transitory computer readable medium of claim 15 wherein the constituent image is a frame of a video.

19. The non-transitory computer readable medium of claim 15 wherein the spatial area is one of a plurality of spatial areas of the input image, and wherein a union of the plurality of spatial areas contains all pixels of the input image.

20. The non-transitory computer readable medium of claim 15 wherein the spatial area is one of a plurality of spatial areas of the input image, and wherein the plurality of spatial areas are uniform in size and shape.

21. The non-transitory computer readable medium of claim 15 wherein the spatial area is one of a plurality of spatial areas of the input image, and wherein at least two of the plurality of spatial areas have different sizes.

22. The non-transitory computer readable medium of claim 15 wherein the spatial area is one of a plurality of spatial areas of the input image, and wherein at least two of the plurality of spatial areas have different shapes.

Assignments (2)
CHANGE OF NAME Recorded Sep 20, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 057552/0227 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2021
From: PAVETIC, FILIP; LEUNG, KING HONG THOMAS; TOCHILKIN, DMITRII
To: GOOGLE INC.
Reel/Frame 057309/0435 →
Continuity (4)
Continuation 16813686 · Mar 9, 2020
Continuation 15444054 · Feb 27, 2017
Provisional Application 62446057 · Jan 13, 2017
Related Publication 20210374418A1 · Dec 2, 2021