IP Library Granted Patent US 10,956,494
Granted Patent B2
US 10,956,494 · App. 16/386,241 · Granted Mar 23, 2021

Behavioral measurements in a video stream focalized on keywords

Inventors: Ramya Narasimha (Palo Alto, CA); Hector H. Gonzalez-Banos (Mountain View, CA)
Assignee: Ricoh Company, Ltd.
G06F16/784G06F16/24578G06F16/71G06F16/738G06F16/739G06F16/7834G06F16/7867G06K9/00288G06K9/00718G06K9/00744G06K9/00758G06K9/00765G06K9/00771G06K9/46G06F16/9566G06K9/00295G06K2009/00738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,956,494
App. No.
16/386,241
Granted
Mar 23, 2021
Kind
B2
Abstract

A system and method for analyzing behavior in a video is described. The method includes extracting a plurality of salient fragments of a video; building a database of the plurality of salient fragments; receiving a keyword; identifying a time anchor when the keyword appears in an audio track associated with the video; retrieving one or more salient fragments of the video from the database of the plurality of salient fragments based on the time anchor; generating a focalized visualization based on the one or more salient fragments of the video; tagging a human subject in the focalized visualization with a unique identifier; analyzing the focalized visualization based on the time anchor and the unique identifier to generate a behavior score; and providing the behavior score via the user device.

Claims (80)

1. A computer-implemented method comprising:

extracting a plurality of salient fragments from a video, a salient fragment being a video sequence tracking a salient object through a subset of a series of frames in the video;

building a database of the plurality of salient fragments;

receiving a keyword;

identifying a time of utterance of the keyword in an audio track of the video;

retrieving one or more salient fragments of the video from the database of the plurality of salient fragments using the time of utterance of the keyword;

generating a video segment focalized on the utterance of the keyword using the one or more retrieved salient fragments of the video, the video segment starting at the time of utterance of the keyword;

tagging a human subject in the video segment with a unique identifier;

analyzing the video segment based on the time of utterance of the keyword and the unique identifier to generate a behavior score; and

providing the behavior score via a user device.

2. The method of claim 1 , further comprising:

storing the behavior score in a behavioral analysis database using the time of utterance of the keyword and the unique identifier as a key to the behavioral analysis database.

3. The method of claim 1 , wherein tagging the human subject in the video segment with the unique identifier comprises:

detecting a face in a frame of the video segment;

identifying the face using a template of a known human subject, wherein the template is associated with the unique identifier; and

associating the face with the unique identifier.

4. The method of claim 1 , further comprising:

tracking the human subject over a plurality of frames of the video segment.

5. The method of claim 1 , wherein analyzing the video segment based on the time of utterance of the keyword and the unique identifier to generate the behavior score comprises:

determining a position of the human subject in a fragment of the video segment at a time corresponding to the time of utterance of the keyword; and

assigning a score to the human subject as a function of the position of the human subject.

6. The method of claim 1 , wherein analyzing the video segment based on the time of utterance of the keyword and the unique identifier to generate the behavior score comprises:

determining a mood of the human subject in a fragment of the video segment at a time corresponding to the time of utterance of the keyword; and

assigning a score to the human subject as a function of the mood of the human subject.

7. The method of claim 1 , further comprising:

generating a baseline score for a behavioral attribute; and

determining a comparison score based on the baseline score and the behavior score.

8. A computer program product comprising a non-transitory computer readable medium storing a computer readable program, wherein the computer readable program when executed on a computer causes the computer to:

extract a plurality of salient fragments from a video, a salient fragment being a video sequence tracking a salient object through a subset of a series of frames in the video;

build a database of the plurality of salient fragments;

receive a keyword;

identify a time of utterance of the keyword in an audio track of the video;

retrieve one or more salient fragments of the video from the database of the plurality of salient fragments using the time of utterance of the keyword;

generate a video segment focalized on the utterance of the keyword using the one or more retrieved salient fragments of the video, the video segment starting at the time of utterance of the keyword;

tag a human subject in the video segment with a unique identifier;

analyze the video segment based on the time of utterance of the keyword and the unique identifier to generate a behavior score; and

provide the behavior score via a user device.

9. The computer program product of claim 8 , wherein the computer readable program when executed on the computer further causes the computer to:

store the behavior score in a behavioral analysis database using the time of utterance of the keyword and the unique identifier as a key to the behavioral analysis database.

10. The computer program product of claim 8 , wherein to tag the human subject in the video segment with the unique identifier, the computer readable program causes the computer to:

detect a face in a frame of the video segment;

identify the face using a template of a known human subject, wherein the template is associated with the unique identifier; and

associate the face with the unique identifier.

11. The computer program product of claim 8 , wherein the computer readable program when executed on the computer further causes the computer to:

track the human subject over a plurality of frames of the video segment.

12. The computer program product of claim 8 , wherein to analyze the video segment based on the time of utterance of the keyword and the unique identifier to generate the behavior score, the computer readable program causes the computer to:

determine a position of the human subject in a fragment of the video segment at a time corresponding to the time of utterance of the keyword; and

assign a score to the human subject as a function of the position of the human subject.

13. The computer program product of claim 8 , wherein to analyze the video segment based on the time of utterance of the keyword and the unique identifier to generate the behavior score, the computer readable program causes the computer to:

determine a mood of the human subject in a fragment of the video segment at a time corresponding to the time of utterance of the keyword; and

assign a score to the human subject as a function of the mood of the human subject.

14. The computer program product of claim 8 , wherein the computer readable program when executed on the computer further causes the computer to:

generate a baseline score for a behavioral attribute; and

determine a comparison score based on the baseline score and the behavior score.

15. A system comprising:

one or more processors; and

a memory, the memory storing instructions which when executed cause the one or more processors to:

extract a plurality of salient fragments from a video, a salient fragment being a video sequence tracking a salient object through a subset of a series of frames in the video;

build a database of the plurality of salient fragments;

receive a keyword;

identify a time of utterance of the keyword in an audio track of the video;

retrieve one or more salient fragments of the video from the database of the plurality of salient fragments using the time of utterance of the keyword;

generate a video segment focalized on the utterance of the keyword using the one or more retrieved salient fragments of the video, the video segment starting at the time of utterance of the keyword;

tag a human subject in the video segment with a unique identifier;

analyze the video segment based on the time of utterance of the keyword and the unique identifier to generate a behavior score; and

provide the behavior score via a user device.

16. The system of claim 15 , wherein the instructions further causes the one or more processors to:

store the behavior score in a behavioral analysis database using the time of utterance of the keyword and the unique identifier as a key to the behavioral analysis database.

17. The system of claim 15 , wherein to tag the human subject in the video segment with the unique identifier, the instructions further causes the one or more processors to:

detect a face in a frame of the video segment;

identify the face using a template of a known human subject, wherein the template is associated with the unique identifier; and

associate the face with the unique identifier.

18. The system of claim 15 , wherein the instructions further causes the one or more processors to:

track the human subject over a plurality of frames of the video segment.

19. The system of claim 15 , wherein to analyze the video segment based on the time of utterance of the keyword and the unique identifier to generate the behavior score, the instructions further causes the one or more processors to:

determine a position of the human subject in a fragment of the video segment at a time corresponding to the time of utterance of the keyword; and

assign a score to the human subject as a function of the position of the human subject.

20. The system of claim 15 , wherein to analyze the video segment based on the time of utterance of the keyword and the unique identifier to generate the behavior score, the instructions further causes the one or more processors to:

determine a mood of the human subject in a fragment of the video segment at a time corresponding to the time of utterance of the keyword; and

assign a score to the human subject as a function of the mood of the human subject.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2019
From: NARASIMHA, RAMYA; GONZALEZ-BANOS, HECTOR H.
To: RICOH COMPANY, LTD.
Reel/Frame 050622/0799 →
Continuity (4)
Continuation In Part 15916997 · Mar 9, 2018
Continuation In Part 15453722 · Mar 8, 2017
Continuation In Part 15447416 · Mar 2, 2017
Related Publication 20190243853A1 · Aug 8, 2019