IP Library Granted Patent US 11,328,178
Granted Patent B2
US 11,328,178 · App. 16/817,364 · Granted May 10, 2022

System and method for automated photo-ideophone matching and placement

Inventors: David Ayman Shamma (San Francisco, CA); Lyndon Kennedy (San Francisco, CA); Anthony Dunnigan (Palo Alto, CA)
Assignee: FUJIFILM Business Innovation Corp.
G06K9/6255G06F40/169G06F40/242G06K9/6267G06V40/161G06V40/172G06V40/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,328,178
App. No.
16/817,364
Granted
May 10, 2022
Kind
B2
Abstract

A computer-implemented method of associating an annotation with an object in an image, comprising generating a dictionary including first vectors that associate terms of the annotation with concepts, classifying the image to generate a second vector based on classified objects and associated confidence scores for the classified objects, selecting a term of the terms associated with one of the first vectors having a shortest determined distance to the second vector, identifying a non-salient region of the image, and rendering the annotation associated with the selected term at the non-salient region.

Claims (30)

1. A computer-implemented method of associating an annotation with an object in an image, comprising:

generating a dictionary including first vectors that associate terms of the annotation with concepts;

classifying the image to generate a second vector based on classified objects and associated confidence scores for the classified objects;

selecting a term of the terms associated with one of the first vectors having a shortest determined distance to the second vector;

identifying a non-salient region of the image; and

rendering the annotation associated with the selected term at the non-salient region.

2. The computer-implemented method of claim 1 , wherein the rendering is formed automatically without requiring an input from a user.

3. The computer-implemented method of claim 1 , wherein the classifying comprises applying one or more of a first classifier associated with detection of a face and classification of an expression of the face, a second classifier associated with classification of food, and a third classifier associated with classification of an item in the image.

4. The computer-implemented method of claim 3 , wherein the classification scores comprise a floating points score between 0 and 1, wherein 0 is indicative of a frown on the face and 1 is indicative of a smile on the face.

5. The computer-implemented method of claim 1 , wherein the determining the non-salient region of the image comprises detecting contours of the image, dividing the image into regions based on its midpoint, and selecting one of the regions having a lowest level of overlap of the contours of the image.

6. The computer-implemented method of claim 1 , wherein determining the shortest determined distance comprises calculation of a cosine distance between the second vector and each of the first vectors.

7. The computer-implemented method of claim 1 , wherein the computer implemented method is executed a standalone device or a cloud connected device as one or more of a video editor, and online photo application and a photo book.

8. The computer-implemented method of claim 1 , wherein the terms comprise one or more of words, languages, fonts and colors, and the one or more of sounds and sentiments.

9. The computer-implemented method of claim 1 , wherein the image comprises a frame of a video, the object is a face, and the annotation is associated with an audio output of the face.

10. The computer-implemented method of claim 1 , wherein the image comprises a frame of a video, the object is an item, and the annotation is associated with an audio output based on a movement of the object.

11. A non-transitory computer readable medium having a storage that stores instructions associating an annotation with an object in an image, the instructions executed by a processor, the instructions comprising:

generating a dictionary including first vectors that associate terms of the annotation with concepts;

classifying the image to generate a second vector based on classified objects and associated confidence scores for the classified objects;

selecting a term of the terms associated with one of the first vectors having a shortest determined distance to the second vector;

identifying a non-salient region of the image; and

rendering the annotation associated with the selected term at the non-salient region.

12. The computer-implemented method of claim 1 , wherein the rendering is formed automatically without requiring an input from a user.

13. The computer-implemented method of claim 1 , wherein the classifying comprises applying one or more of a first classifier associated with detection of a face and classification of an expression of the face, a second classifier associated with classification of food, and a third classifier associated with classification of an item in the image.

14. The computer-implemented method of claim 3 , wherein the classification scores comprise a floating points score between 0 and 1, wherein 0 is indicative of a frown on the face and 1 is indicative of a smile on the face.

15. The computer-implemented method of claim 1 , wherein the determining the non-salient region of the image comprises detecting contours of the image, dividing the image into regions based on its midpoint, and selecting one of the regions having a lowest level of overlap of the contours of the image.

16. The computer-implemented method of claim 1 , wherein determining the shortest determined distance comprises calculation of a cosine distance between the second vector and each of the first vectors.

17. The computer-implemented method of claim 1 , wherein the computer implemented method is executed a standalone device or a cloud connected device as one or more of a video editor, and online photo application and a photo book.

18. The computer-implemented method of claim 1 , wherein the terms comprise one or more of words, languages, fonts and colors, and the one or more of sounds and sentiments.

19. The computer-implemented method of claim 1 , wherein the image comprises a frame of a video, the object is a face, and the annotation is associated with an audio output of the face.

20. The computer-implemented method of claim 1 , wherein the image comprises a frame of a video, the object is an item, and the annotation is associated with an audio output based on a movement of the object.

Assignments (2)
CHANGE OF NAME Recorded May 25, 2021
From: FUJI XEROX CO., LTD.
To: FUJIFILM BUSINESS INNOVATION CORP.
Reel/Frame 056392/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2020
From: SHAMMA, DAVID AYMAN; KENNEDY, LYNDON; DUNNIGAN, ANTHONY
To: FUJI XEROX CO., LTD.
Reel/Frame 052102/0116 →
Continuity (1)
Related Publication 20210287039A1 · Sep 16, 2021