IP Library Granted Patent US 12,217,142
Granted Patent B2
US 12,217,142 · App. 18/520,532 · Granted Feb 4, 2025

Using machine learning to detect which part of the screen includes embedded frames of an uploaded video

Inventors: Filip Pavetic (Zürich, CH); King Hong Thomas Leung (Saratoga, CA); Dmitrii Tochilkin (Zürich, CH)
Assignee: Google LLC
G06N20/00G06F18/214G06F18/241G06V10/25G06V10/764G06V20/40G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,142
App. No.
18/520,532
Granted
Feb 4, 2025
Kind
B2
Abstract

A system and methods are disclosed for using a trained machine learning model to identify constituent images within composite images. A method may include providing data identifying a first image as input to a machine learning model trained using training data identifying a plurality of composite images that each include one or more constituent images, and determining, using one or more outputs of the trained machine learning model, that the first image is a composite image that includes a first constituent image, wherein at least a portion of the first constituent image is in a spatial area of the first image, and wherein the first constituent image corresponds to a frame of a video embedded into the first image.

Claims (45)

1. A method comprising:

providing data identifying a first image as input to a machine learning model trained using training data identifying a plurality of composite images that each include one or more constituent images; and

determining, using one or more outputs of the trained machine learning model, that the first image is a composite image that includes a first constituent image, wherein at least a portion of the first constituent image is in a spatial area of the first image, and wherein the first constituent image corresponds to a frame of a video embedded into the first image.

2. The method of claim 1 , wherein:

the one or more outputs of the trained machine learning model comprise a level of confidence value indicating a likelihood of the first image being the composite image including the first constituent image; and

the method further comprises:

determining that the level of confidence value satisfies a threshold condition; and

extracting the first constituent image from the spatial area of the first image.

3. The method of claim 1 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein the plurality of spatial areas are uniform in size and shape.

4. The method of claim 1 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas have different sizes.

5. The method of claim 1 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas overlap.

6. The method of claim 1 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein the plurality of spatial areas are non-overlapping.

7. The method of claim 1 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas have different shapes.

8. The method of claim 1 , wherein the training data comprises an input-output mapping comprising an input and an output, the input based on pixel data of a composite image, the composite image comprising a first portion containing pixel data of a second image and a second portion containing pixel data of a third image, and the output identifying a position of the first portion within the composite image.

9. The method of claim 8 wherein the position of the first portion within the composite image comprises coordinates of an upper left corner of the second image and coordinates of a lower right corner of the second image.

10. A system comprising:

a memory; and

a processing device, coupled to the memory, to perform operations comprising:

providing data identifying a first image as input to a machine learning model trained using training data identifying a plurality of composite images that each include one or more constituent images; and

determining, using one or more outputs of the trained machine learning model, that the first image is a composite image that includes a first constituent image, wherein at least a portion of the first constituent image is in a spatial area of the first image, and wherein the first constituent image corresponds to a frame of a video embedded into the first image.

11. The system of claim 10 , wherein:

the one or more outputs of the trained machine learning model comprise a level of confidence value indicating a likelihood of the first image being the composite image including the first constituent image; and

the operations further comprise:

determining that the level of confidence value satisfies a threshold condition; and

extracting the first constituent image from the spatial area of the first image.

12. The system of claim 10 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein the plurality of spatial areas are uniform in size and shape.

13. The system of claim 10 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas have different sizes.

14. The system of claim 10 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas overlap.

15. The system of claim 10 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein the plurality of spatial areas are non-overlapping.

16. The system of claim 10 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein at least two of the plurality of spatial areas have different shapes.

17. The system of claim 10 , wherein the training data comprises an input-output mapping comprising an input and an output, the input based on pixel data of a composite image, the composite image comprising a first portion containing pixel data of a second image and a second portion containing pixel data of a third image, and the output identifying a position of the first portion within the composite image.

18. A non-transitory computer readable medium comprising instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

providing data identifying a first image as input to a machine learning model trained using training data identifying a plurality of composite images that each include one or more constituent images; and

determining, using one or more outputs of the trained machine learning model, that the first image is a composite image that includes a first constituent image, wherein at least a portion of the first constituent image is in a spatial area of the first image, and wherein the first constituent image corresponds to a frame of a video embedded into the first image.

19. The non-transitory computer readable medium of claim 18 , wherein:

the one or more outputs of the trained machine learning model comprise a level of confidence value indicating a likelihood of the first image being the composite image including the first constituent image; and

the operations further comprise:

determining that the level of confidence value satisfies a threshold condition; and

extracting the first constituent image from the spatial area of the first image.

20. The non-transitory computer readable medium of claim 18 , wherein the spatial area is one of a plurality of spatial areas of the first image, and wherein one or more of the following conditions are satisfied:

the plurality of spatial areas are uniform in size and shape;

at least two of the plurality of spatial areas have different sizes;

at least two of the plurality of spatial areas overlap;

the plurality of spatial areas are non-overlapping; or

at least two of the plurality of spatial areas have different shapes.

Assignments (2)
CHANGE OF NAME Recorded Feb 1, 2024
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 066488/0574 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2024
From: PAVETIC, FILIP; LEUNG, KING HONG THOMAS; TOCHILKIN, DMITRII
To: GOOGLE INC.
Reel/Frame 066187/0091 →
Continuity (5)
Continuation 17403804 · Aug 16, 2021
Continuation 16813686 · Mar 9, 2020
Continuation 15444054 · Feb 27, 2017
Provisional Application 62446057 · Jan 13, 2017
Related Publication 20240104435A1 · Mar 28, 2024
References Cited (44)
US 1020427A · Kellogg · 1912 [cited by applicant]
US 7697024B2 · Currivan et al. · 2010 [cited by applicant]
US 8447708B2 · Sabe · 2013 [cited by applicant]
US 8761498B1 · Wu · 2014 [cited by applicant]
US 9129148B1 · Li · 2015 [cited by examiner]
US 9148699B2 · Shivalingappa et al. · 2015 [cited by applicant]
US 9357117B2 · Woo et al. · 2016 [cited by applicant]
US 9787938B2 · Cranfill et al. · 2017 [cited by applicant]
US 9973722B2 · Deng et al. · 2018 [cited by applicant]
US 9992410B2 · Kim et al. · 2018 [cited by applicant]
US 10157332B1 · Gray · 2018 [cited by examiner]
US 10204274B2 · Smith, IV et al. · 2019 [cited by applicant]
US 10230866B1 · Townsend et al. · 2019 [cited by applicant]
US 10346723B2 · Han · 2019 [cited by examiner]
US 10521671B2 · Chang · 2019 [cited by examiner]
US 10580179B2 · Luan et al. · 2020 [cited by applicant]
US 10984572B1 · Zacharia et al. · 2021 [cited by applicant]
US 11107232B2 · Li · 2021 [cited by applicant]
US 20090208118A1 · Csurka · 2009 [cited by applicant]
US 20100067865A1 · Saxena et al. · 2010 [cited by applicant]
US 20140363143A1 · Dharssi et al. · 2014 [cited by applicant]
US 20180121392A1 · Zhang et al. · 2018 [cited by applicant]
US 20180121732A1 · Kim et al. · 2018 [cited by applicant]
US 20180121762A1 · Han et al. · 2018 [cited by applicant]
US 20180204065A1 · Pavetic et al. · 2018 [cited by applicant]
US 20180260668A1 · Shen et al. · 2018 [cited by applicant]
US 20200210709A1 · Pavetic et al. · 2020 [cited by applicant]
CN 103679142A · 2014 [cited by applicant]
CN 105678338A · 2016 [cited by applicant]
US 10,019,463 B2, 07/2018, Li et al. (withdrawn) [cited by applicant]
Constine J., “Facebook Launches Video Rights Manager to Combat Freebooting,” Report of F8 Facebook Developer Conference held on Apr. 12-13, 2016, on TC's Crunchboard, 7 Pages, [Retrieved on Jan. 13, 2017], Retrieved fro… [cited by applicant]
Digital Video Fingerprinting. In: Wikipedia, The free encyclopedia. Edit date: Dec. 29, 2016, 18:36 UTC. URL: https://en.wikipedia.org/w/index.php?title=Digital_video_fingerprintingoldid=757261679 [accessed Aug. 3, 2023… [cited by applicant]
Erhan D., et al., “Scalable Object Detection using Deep Neural Networks,” 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 23-28, 2014, 8 Pages. [cited by applicant]
Intellectual Property Office, Combined Search and Examination Report under Sections 17 and 18(3), Office Action for Application No. 1717849.2, mailed Apr. 30, 2018, 6 Pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2017/057668, mailed Feb. 6, 2018, 13 Pages. [cited by applicant]
Office Action for German Patent Application No. DE201710125463, mailed Sep. 8, 2023, 16 Pages. [cited by applicant]
Pratusevich M., “EdVidParse: Detecting People and Content in Educational Videos,” Massachusetts Institute of Technology, 2015, 65 Pages. [cited by applicant]
Redmon J., et al., “YOLO9000: Better, Faster, Stronger,” University of Washington, Allen Institute for AI, arXiv:1612.08242v1, Dec. 25, 2016, 9 Pages, Retrieved from URL: http://pjreddie.com/yolo9000. [cited by applicant]
Rozantsev A., “On Rendering Synthetic Images for Training an Object Detector,” Ecole Polytechnique Federale de Lausanne, Computer Vision Laboratory, Lausanne, Switzerland, Graz University of Technology, Institute for Co… [cited by applicant]
Szegedy C., et al., “Deep Neural Networks for Object Detection,” Proceedings of the 26th International Conference on Advances in Neural Information Processing Systems (NPIS), Dec. 5-10, 2013, 9 Pages. [cited by applicant]
Taschwer M., et al., “Compound Figure Separation Combining Edge and Band Separator Detection,” ITEC, Klagenfurt University (MU), Klagenfurt, Austria; Florida Atlantic University (FAU), Boca Raton, FL, USA, Springer Inte… [cited by applicant]
Tsutsui S., et al., “A Data Driven Approach for Compound Figure Separation Using Convolutional Neural Networks,” School of Informatics and Computing, Indiana University, Bloomington, Indiana, USA, Aug. 21, 2017, 8 Pages. [cited by applicant]
Written Opinion for International Application No. PCT/US2017/057668, mailed Feb. 6, 2018, 10 Pages. [cited by applicant]
Yu J., et al., “Improving Person Detection Using Synthetic Training data,” Proceedings of 2010 17th IEEE International Conference on Image Processing, Hong Kong, China, Ieee, Piscataway, NJ, USA, Sep. 26-29, 2010, pp. 3… [cited by applicant]