IP Library Granted Patent US 12,288,371
Granted Patent B2
US 12,288,371 · App. 17/884,607 · Granted Apr 29, 2025

Finding the semantic region of interest in images

Inventor: Robert Gonsalves (Wellesley, MA)
Assignee: AVID TECHNOLOGY, INC.
G06V10/25G06T5/94G06T7/11G06T7/70G06T9/00G06T2207/10016G06T2207/20081G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,371
App. No.
17/884,607
Granted
Apr 29, 2025
Kind
B2
Abstract

Objects within an image are assigned a semantic interest that is indicative of their importance to the image as a whole. The objects are detected automatically, and sub-images that span each of the objects are excerpted from the image. An image embedding is determined for each of the sub-images as well as for the whole image by using an image encoder implemented as a trained multi-modal neural network. The degree of similarity between the image embeddings of a sub-image and that of the whole image is used as a measure of the semantic importance of the object portrayed in the sub-image. Objects of high semantic importance comprise the semantic regions of interest of the image. Knowledge of such regions may be used to enhance and improve the efficiency of downstream image-processing tasks, such as image compression, pan and scan, and contrast enhancement.

Claims (56)

1. A method of determining semantic regions of interest within a source image, the method comprising:

receiving the source image;

using an automatic object-detection system to detect a plurality of objects within the source image;

subdividing the source image into a plurality of sub-images, each sub-image containing a portion of the source image that contains one of the detected plurality of objects;

using a trained neural network model to:

generate an image embedding for the source image; and

for each sub-image of the plurality of sub-images, generate an image embedding for the sub-image; and

for each sub-image of the plurality of sub-images:

determining a degree of similarity between the image embedding of the sub-image and the image embedding of the source image;

assigning a semantic interest to the detected object contained by the sub-image according to the determined degree of similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image; and

outputting an indication of the semantic interest assigned to the detected object contained by the sub-image.

2. The method of claim 1 , wherein the automatic object-detection system is a trained neural-network model.

3. The method of claim 1 , wherein the trained neural network model that is used to generate the image embeddings is a multi-modal neural network.

4. The method of claim 1 , further comprising:

for each detected object of the plurality of detected objects, generating an object mask for the detected object;

generating an object mask image of the source image in which each detected object of the plurality of detected objects is replaced in the source image with a shaded silhouette of the object mask generated for the object; and

applying a visual indication to each shaded silhouette, wherein the visual indication is indicative of the semantic interest assigned to the detected object corresponding the object mask.

5. The method of claim 1 , wherein the indication of the semantic interest assigned to each detected object of the plurality of detected objects is used to enhance image-processing of the source image.

6. The method of claim 5 , wherein the image processing comprises image compression, and enhancing the image processing of the source image includes varying a number of bits allocated to compressing each sub-image of the plurality of sub-images in accordance with the semantic interest assigned to the detected object corresponding to the sub-image.

7. The method of claim 6 , wherein the source image is a frame of a video stream.

8. The method of claim 5 , wherein:

the image processing includes cropping a portion of the source image in order to achieve a desired aspect ratio of the source image; and

enhancing the image processing of the source image includes preferentially retaining within the cropped portion of the source image objects to which higher semantic interest have been assigned.

9. The method of claim 8 , wherein the objects that are preferentially retained within the cropped image include an object to which a maximum semantic interest has been assigned.

10. The method of claim 8 , further comprising:

selecting a subset of detected objects of the plurality of detected objects, wherein the selected subset of objects includes a set of objects to which high semantic interest has been assigned;

locating a centroid of the subset of the detected objects within the source image; and

cropping the source image such that the centroid of the subset of the detected objects within the source image is located at a center of the cropped image.

11. The method of claim 8 , wherein the source image is a frame of a video stream.

12. The method of claim 5 , wherein the image processing includes contrast enhancement, and the contrast enhancement includes boosting contrast in a region of the source image containing a detected object to which a high semantic interest has been assigned.

13. The method of claim 12 , wherein the source image is a frame of a video stream.

14. A computer program product comprising:

a non-transitory computer-readable medium with computer-readable instructions encoded thereon, wherein the computer-readable instructions, when processed by a processing device, instruct the processing device to perform a method of determining semantic regions of interest within a source image, the method comprising:

receiving the source image;

using an automatic object-detection system to detect a plurality of objects within the source image;

subdividing the source image into a plurality of sub-images, each sub-image containing a portion of the source image that contains one of the detected plurality of objects;

using a trained neural network model to:

generate an image embedding for the source image; and

for each sub-image of the plurality of sub-images, generate an image embedding for the sub-image; and

for each sub-image of the plurality of sub-images:

determining a degree of similarity between the image embedding of the sub-image and the image embedding of the source image;

assigning a semantic interest to the detected object contained by the sub-image according to the determined degree of similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image; and

outputting an indication of the semantic interest assigned to the detected object contained by the sub-image.

15. A system comprising:

a memory for storing computer-readable instructions; and

a processor connected to the memory, wherein the processor, when executing the computer-readable instructions, causes the system to perform a method of determining semantic regions of interest within a source image, the method comprising:

receiving the source image;

using an automatic object-detection system to detect a plurality of objects within the source image;

subdividing the source image into a plurality of sub-images, each sub-image containing a portion of the source image that contains one of the detected plurality of objects;

using a trained neural network model to:

generate an image embedding for the source image; and

for each sub-image of the plurality of sub-images, generate an image embedding for the sub-image; and

for each sub-image of the plurality of sub-images:

determining a degree of similarity between the image embedding of the sub-image and the image embedding of the source image;

assigning a semantic interest to the detected object contained by the sub-image according to the determined degree of similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image; and

outputting an indication of the semantic interest assigned to the detected object contained by the sub-image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2022
From: GONSALVES, ROBERT A.
To: AVID TECHNOLOGY, INC.
Reel/Frame 061220/0986 →
Continuity (1)
Related Publication 20240054748A1 · Feb 15, 2024
References Cited (30)
US 20100091330A1 · Marchesotti · 2010 [cited by examiner]
US 20180365835A1 · Yan · 2018 [cited by examiner]
US 20190050639A1 · Ast · 2019 [cited by examiner]
US 20190050648A1 · Stojanovic · 2019 [cited by examiner]
US 20190065911A1 · Lee · 2019 [cited by examiner]
US 20190080204A1 · Schroff · 2019 [cited by examiner]
US 20190286930A1 · Han · 2019 [cited by examiner]
US 20200134375A1 · Zhan · 2020 [cited by examiner]
US 20210248748A1 · Turgutlu · 2021 [cited by examiner]
US 20210264226A1 · Lecue · 2021 [cited by examiner]
US 20220147743A1 · Roy · 2022 [cited by examiner]
US 20220292685A1 · Heisler · 2022 [cited by examiner]
US 20220319141A1 · Liu · 2022 [cited by examiner]
US 20230040513A1 · Ryan · 2023 [cited by examiner]
US 20230386052A1 · Lyu · 2023 [cited by examiner]
CN 114037640A · 2022 [cited by examiner]
EP 3869784 · 2021 [cited by applicant]
JP 202068521 · 2020 [cited by applicant]
He, Mask R-CNN, arXiv: 1703.0687043, Jan. 24, 2018, 12 pgs. [cited by applicant]
Li, Mask Dino: Towards A Unified Transformer-Based Framework for Object Detection and Segmentation, arXiv: 2206.0277v1, Jun. 6, 2022, 13 pages. [cited by applicant]
Li, Oscar: Object Semantics Aligned Pre-training for Vision-Language Tasks, arXiv: 2004.06165v5, Jul. 26, 2020, 21 pages. [cited by applicant]
Radford, Learning Transferable Visual Modes from Natural Language Supervision, arXvi: 2018.00020v1, Feb. 26, 2021, 48 pages. [cited by applicant]
Rippel, ELF-VC: Efficient Learned Flexible-Rate Video Coding, 2021 IEEE/CVF International Conference on Computer Vision, Apr. 2021, 14 pages. [cited by applicant]
Rossi, A Novel Region of Interest Extraction Layer for Instance Segmentation, arXiv: 2004.13665v2, Oct. 1, 2020, 7 pages. [cited by applicant]
Salton, Term Weighting Approaches in Automatic Text Retrieval, Dept. of Computer Science, Cornell University, #87-881, Nov. 1987, 23 pages. [cited by applicant]
Wallace, The JPEG Still Picture Compression Standard, submitted to IEEE Transactions on Consumer Electronics, Dec. 1991, 17 pages. [cited by applicant]
Zhai, LiT: Zero-Shot Transfers with Locked-image Tuning, arXiv: 2111.07991v3, Jun. 22, 2022, 26 pages. [cited by applicant]
Caron et al: “Use of power law models in detecting region of interest”, Pattern Recognition, Elsevier, GB, vol. 40, No. 9, May 3, 2007 (May 3, 2007), pp. 2521-2529, XP022058965, ISSN: 0031-3203, DOI: 10.1016/J.PATCOG.20… [cited by applicant]
Umeki Yo et al: “Salient Object Detection With Importance Degree”, IEEE Access, IEEE, USA, vol. 8, Aug. 6, 2020 (Aug. 6, 2020), pp. 147059-147069, XP011805960, DOI: 10.1109/ACCESS.2020.3014886. [cited by applicant]
Wang Baoyan et al: “Salient object detection based on objectness”, 2015 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), IEEE, Sep. 19, 2015 (Sep. 19, 2015), pp. 1-5, XP03281777… [cited by applicant]