IP Library › Granted Patent US 11,947,631
Granted Patent B2
US 11,947,631 · App. 17/482,290 · Granted Apr 2, 2024

Reverse image search based on deep neural network (DNN) model and image-feature detection model

Inventors: Jong Hwa Lee (San Diego, CA); Praggya Garg (Rancho Santa Fe, CA)
Assignee: SONY GROUP CORPORATION
G06F18/22G06F3/14G06F18/2431G06F18/253G06F18/40G06N3/04G06N3/08G06T7/0002G06V10/40G06V20/46G06T2207/10016G06T2207/20084G06T2207/30168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,947,631
App. No.
17/482,290
Granted
Apr 2, 2024
Kind
B2
Abstract

An electronic device and method for reverse image search is provided. The electronic device receives an image. The electronic device extracts, by a DNN model, a first set of image features associated with the image and generates a first feature vector based on the first set of image features. The electronic device extracts, by an image-feature detection model, a second set of image features associated with the image and generates a second feature vector based on the second set of image features. The electronic device generates a third feature vector based on combination of the first and second feature vectors. The electronic device determines a similarity metric between the third feature vector and a fourth feature vector of each of a set of pre-stored images and identifies a pre-stored image based on the similarity metric. The electronic device controls a display device to display information associated with the pre-stored image.

Claims (85)

1. An electronic device, comprising:

circuitry configured to:

receive a first image;

extract, by a Deep Neural Network (DNN) model, a first set of image features associated with the received first image;

generate a first feature vector associated with the received first image based on the extracted first set of image features;

extract, by an image-feature detection model, a second set of image features associated with the received first image;

generate a second feature vector associated with the received first image based on the extracted second set of image features;

determine, based on a resolution of the received first image, each of:

a first weight associated with the generated first feature vector, and

a second weight associated with the generated second feature vector;

combine the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight;

generate a third feature vector associated with the received first image based on the combination of the generated first feature vector and the generated second feature vector;

determine a similarity metric between the generated third feature vector associated with the received first image and a fourth feature vector of each of a set of pre-stored second images;

identify a pre-stored third image from the set of pre-stored second images based on the determined similarity metric; and

control a display device to display information associated with the identified pre-stored third image.

2. The electronic device according to claim 1 , wherein the image-feature detection model comprises at least one of a Scale-Invariant Feature Transform (SIFT)-based model, a Speeded-Up Robust Features (SURF)-based model, an Oriented FAST and Rotated BRIEF (ORB)-based model, or a Fast Library for Approximate Nearest Neighbors (FLANN)-based model.

3. The electronic device according to claim 1 , wherein the generation of the third feature vector is further based on an application of a Principal Component Analysis (PCA) transformation on the combination of the generated first feature vector and the generated second feature vector.

4. The electronic device according to claim 1 , wherein the similarity metric comprises at least one of a cosine-distance similarity or a Euclidean-distance similarity.

5. The electronic device according to claim 1 , wherein the circuitry is further configured to determine, by a machine learning model different from the DNN model and the image-feature detection model, the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector.

6. The electronic device according to claim 1 , wherein the circuitry is further configured to:

receive a user input including the first weight associated with the generated first feature vector and including the second weight associated with the generated second feature vector;

combine the generated first feature vector and the generated second feature vector based on the received user input; and

generate the third feature vector based on the combination of the generated first feature vector and the generated second feature vector.

7. The electronic device according to claim 1 , wherein the circuitry is further configured to:

classify, by the DNN model, the received first image into a first image tag from a set of image tags associated with the DNN model;

determine a first count of images, associated with the first image tag, in a training dataset associated with the DNN model;

determine the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector based on the determined first count of images associated with the first image tag;

combine the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight; and

generate the third feature vector based on the combination of the generated first feature vector and the generated second feature vector.

8. The electronic device according to claim 1 , wherein the circuitry is further configured to:

determine an image quality score associated with the received first image based on at least one of the extracted first set of image features or the extracted second set of image features;

determine the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector based on the determined image quality score;

combine the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight; and

generate the third feature vector based on the combination of the generated first feature vector and the generated second feature vector.

9. The electronic device according to claim 8 , wherein the image quality score corresponds to at least one of a sharpness, a noise, a dynamic range, a tone reproduction, a contrast, a color saturation, a distortion, a vignetting, an exposure accuracy, a chromatic aberration, a lens flare, a color moire, or artifacts associated with the received first image.

10. The electronic device according to claim 1 , wherein the circuitry is further configured to extract the first image from a first video and the identified pre-stored third image that corresponds to a pre-stored second video, and wherein the pre-stored second video is associated with the first video.

11. A method, comprising:

in an electronic device:

receiving a first image;

extracting, by a Deep Neural Network (DNN) model, a first set of image features associated with the received first image;

generating a first feature vector associated with the received first image based on the extracted first set of image features;

extracting, by an image-feature detection model, a second set of image features associated with the received first image;

generating a second feature vector associated with the received first image based on the extracted second set of image features;

determining, based on a resolution of the received first image, each of:

a first weight associated with the generated first feature vector, and

a second weight associated with the generated second feature vector;

combining the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight;

generating a third feature vector associated with the received first image based on the combination of the generated first feature vector and the generated second feature vector;

determining a similarity metric between the generated third feature vector associated with the received first image and a fourth feature vector of each of a set of pre-stored second images;

identifying a pre-stored third image from the set of pre-stored second images based on the determined similarity metric; and

controlling a display device to display information associated with the identified pre-stored third image.

12. The method according to claim 11 , wherein the image-feature detection model comprises at least one of a Scale-Invariant Feature Transform (SIFT)-based model, a Speeded-Up Robust Features (SURF)-based model, an Oriented FAST and Rotated BRIEF (ORB)-based model, or a Fast Library for Approximate Nearest Neighbors (FLANN)-based model.

13. The method according to claim 11 , wherein the generation of the third feature vector is further based on an application of a Principal Component Analysis (PCA) transformation on the combination of the generated first feature vector and the generated second feature vector.

14. The method according to claim 11 , wherein the similarity metric comprises at least one of a cosine-distance similarity or a Euclidean-distance similarity.

15. The method according to claim 11 , further comprising determining, by a machine learning model different from the DNN model and the image-feature detection model, the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector.

16. The method according to claim 11 , further comprising:

receiving a user input including the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector;

combining the generated first feature vector and the generated second feature vector based on the received user input; and

generating the third feature vector based on the combination of the generated first feature vector and the generated second feature vector.

17. The method according to claim 11 , further comprising:

classifying, by the DNN model, the received first image into a first image tag from a set of image tags associated with the DNN model;

determining a first count of images, associated with the first image tag, in a training dataset associated with the DNN model;

determining the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector based on the determined first count of images associated with the first image tag;

combining the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight; and

generating the third feature vector based on the combination of the generated first feature vector and the generated second feature vector.

18. The method according to claim 11 , further comprising:

determining an image quality score associated with the received first image based on at least one of the extracted first set of image features or the extracted second set of image features;

determining the first weight associated with the generated first feature vector and the second weight associated with the generated second feature vector based on the determined image quality score;

combining the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight; and

generate the third feature vector based on the combination of the generated first feature vector and the generated second feature vector.

19. The method according to claim 18 , wherein the image quality score corresponds to at least one of a sharpness, a noise, a dynamic range, a tone reproduction, a contrast, a color saturation, a distortion, a vignetting, an exposure accuracy, a chromatic aberration, a lens flare, a color moire, or artifacts associated with the received first image.

20. A non-transitory computer-readable medium having stored thereon, computer-executable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprising:

receiving a first image;

extracting, by a Deep Neural Network (DNN) model, a first set of image features associated with the received first image;

generating a first feature vector associated with the received first image based on the extracted first set of image features;

extracting, by an image-feature detection model, a second set of image features associated with the received first image;

generating a second feature vector associated with the received first image based on the extracted second set of image features;

determining, based on a resolution of the received first image, each of:

a first weight associated with the generated first feature vector, and

a second weight associated with the generated second feature vector;

combining the generated first feature vector and the generated second feature vector based on the determined first weight and the determined second weight;

generating a third feature vector associated with the received first image based on the combination of the generated first feature vector and the generated second feature vector;

determining a similarity metric between the generated third feature vector associated with the received first image and a fourth feature vector of each of a set of pre-stored second images;

identifying a pre-stored third image from the set of pre-stored second images based on the determined similarity metric; and

controlling a display device to display information associated with the identified pre-stored third image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2021
From: LEE, JONG HWA; GARG, PRAGGYA
To: SONY GROUP CORPORATION
Reel/Frame 058506/0252 →
Continuity (2)
Provisional Application 63189956 · May 18, 2021
Related Publication 20220374647A1 · Nov 24, 2022