IP Library Granted Patent US 11,481,575
Granted Patent B2
US 11,481,575 · App. 16/142,155 · Granted Oct 25, 2022

System and method for learning scene embeddings via visual semantics and application thereof

Inventors: Paloma de Juan (New York, NY); Aasish Pappu (New York, NY)
Assignee: YAHOO ASSETS LLC
G06K9/6255G06F16/5838G06K9/6262G06N3/08G06N20/00G06V30/274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,575
App. No.
16/142,155
Granted
Oct 25, 2022
Kind
B2
Abstract

The present teaching relates to method, system, and programming for responding to an image related query. Information related to each of a plurality of images is received, wherein the information represents concepts co-existing in the image. Visual semantics for each of the plurality of images are created based on the information related thereto. Representations of scenes of the plurality of images are obtained via machine learning, based on the visual semantics of the plurality of images, wherein the representations capture concepts associated with the scenes.

Claims (55)

1. A method, implemented on a machine having at least one processor, storage, and a communication platform for responding to an image related query, comprising:

receiving, via the communication platform, information related to each image of a plurality of images, wherein the information indicates concepts co-existing in each image of the plurality of images;

creating visual semantics for each image of the plurality of images based on the information related thereto, wherein the visual semantics for each image of the plurality of images comprises a hierarchy of abstraction having multiple abstraction levels, including a top level of abstraction on the nature of an entire scene of the image, an intermediate level of abstraction on each category of each object in the image, and a low level of abstraction on each instance of a corresponding object category corresponding to each object in the image; and

obtaining via machine learning, representation of image scenes based on the visual semantics of the plurality of images, wherein the representation of each image scene conceptually summarizes the scene based on spatial relationships among concepts co-existing in the image.

2. The method of claim 1 , wherein the visual semantics created for each of the plurality of images includes an identifier of each image of the plurality of images and one or more annotations of the concepts co-existing in each image of the plurality of images.

3. The method of claim 2 , wherein

the identifier of each image of the plurality of images provides a context of the visual semantics; and

the one or more annotations specify the concepts co-existing in each image of the plurality of images.

4. The method of claim 2 , wherein the representation of the corresponding scene includes:

a plurality of vectors for the one or more annotations related to the concepts co-existing in each image of the plurality of images, the identifier of each image of the plurality of images, and at least one combination thereof; and

an artificial neural network (ANN) with a plurality of layers of nodes and connections therein connecting the nodes.

5. The method of claim 1 , wherein the representation of the corresponding scene corresponds to a scene embedding.

6. The method of claim 5 , wherein the scene embedding corresponds to hierarchical relationships between the concepts co-existing in each image of the

plurality of images associated with the scene.

7. The method of claim 1 , further comprising:

receiving the image related query;

obtaining a response to the image related query based on representations obtained via the machine learning.

8. The method of claim 7 , wherein the image related query is directed to at least one of:

a request to receive a summary of at least one concept, from the concepts co-existing in each of the plurality of images, included in the image related query; and

a request to receive one or more images that meet a conceptual similarity criterion with respect to the image related query.

9. A system for responding to an image related query, the system comprising:

a visual semantics generator implemented by a processor and configured to

receive information related to each image of a plurality of images, wherein the information indicates concepts co-existing in each image of the plurality of images, and

create visual semantics for each image of the plurality of images based on the information related thereto, wherein the visual semantics for each image of the plurality of images comprises a hierarchy of abstraction having multiple abstraction levels, including a top level of abstraction on the nature of an entire scene of the image, an intermediate level of abstraction on each category of each object in the image, and a low level of abstraction on each instance of a corresponding object category corresponding to each object in the image; and

an image scene embedding training unit implemented by the processor and configured to obtain, via machine learning, representation of a image scenes based on the visual semantics of the plurality of images, wherein the representation of each image scene conceptually summarizes the scene based on spatial relationships among concepts co-existing in the image.

10. The system of claim 9 , wherein the visual semantics created for each of the plurality of images includes an identifier of each image of the plurality of images and one or more annotations of the concepts co-existing in each image of the plurality of images.

11. The system of claim 10 , wherein

the identifier of each image of the plurality of images provides a context of the visual semantics; and

the one or more annotations specify the concepts co-existing in each image of the plurality of images.

12. The system of claim 10 , wherein the representation of the corresponding scene includes:

a plurality of vectors for the one or more annotations related to the concepts co-existing in each image of the plurality of images, the identifier of each image of the plurality of images, and at least one combination thereof; and

an artificial neural network (ANN) with a plurality of layers of nodes and connections therein connecting the nodes.

13. The system of claim 9 , wherein the representation of the corresponding scene corresponds to a scene embedding.

14. The system of claim 9 , further comprising:

a visual scene based query engine implemented by the processor and configured to

receive the image related query, and

obtain a response to the image related query based on representations obtained via the machine learning.

15. A machine readable and non-transitory medium having information including machine executable instructions stored thereon for responding to an image related query, wherein the information, when read by a machine, causes the machine to perform:

receiving information related to each image of a plurality of images, wherein the information indicates concepts co-existing in each image of the plurality of images;

creating visual semantics for each image of the plurality of images based on the information related thereto, wherein the visual semantics for each image of the plurality of images comprises a hierarchy of abstraction having multiple abstraction levels, including a top level of abstraction on the nature of an entire scene of the image, an intermediate level of abstraction on each category of each object in the image, and a low level of abstraction on each instance of a corresponding object category corresponding to each object in the image; and

obtaining, via machine learning representation of image scenes based on the visual semantics of the plurality of images, wherein the representation of each image scene conceptually summarizes the scene based on spatial relationships among concepts co-existing in the image.

16. The medium of claim 15 , wherein the visual semantics created for each of the plurality of images includes an identifier of each image of the plurality of images and one or more annotations of the concepts co-existing in each image of the plurality of images.

17. The medium of claim 16 , wherein

the identifier of each image of the plurality of images provides a context of the visual semantics; and

the one or more annotations specify the concepts co-existing in each image of the plurality of images.

18. The medium of claim 16 , wherein the representation of the corresponding scene includes:

a plurality of vectors for the one or more annotations related to the concepts co-existing in each image of the plurality of images, the identifier of each image of the plurality of images, and at least one combination thereof; and

an artificial neural network (ANN) with a plurality of layers of nodes and connections therein connecting the nodes.

19. The medium of claim 15 , wherein the representation of the corresponding scene corresponds to a scene embedding.

20. The medium of claim 15 , wherein the information, when read by the machine, further causes the machine to perform:

receiving the image related query;

obtaining a response to the image related query based on representations obtained via the machine learning.

21. The medium of claim 20 , wherein the image related query is directed to at least one of:

a request to receive a summary of at least one concept, from the concepts co-existing in each of the plurality of images, included in the image related query; and

a request to receive one or more images that meet a conceptual similarity criterion with respect to the image related query.

Assignments (4)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2018
From: DE JUAN, PALOMA; PAPPU, AASISH
To: OATH INC.
Reel/Frame 047638/0085 →