IP Library Granted Patent US 11,048,919
Granted Patent B1
US 11,048,919 · App. 15/993,222 · Granted Jun 29, 2021

Person tracking across video instances

Inventors: Davide Modolo (Seattle, WA); Hao Chen (Kirkland, WA); Enrica Maria Filippi (Seattle, WA); Stephen Gould (Seattle, WA); Camille Claire Le Men (Seattle, WA); Andrea Olgiati (Gilroy, CA)
Assignee: Amazon Technologies, Inc.
G06K9/00295G06K9/00369G06K9/00765G06K9/00771H04N7/181
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,048,919
App. No.
15/993,222
Granted
Jun 29, 2021
Kind
B1
Abstract

People can be tracked across multiple segments of video data, which can correspond to different scenes in a single video file, or multiple video streams or feeds. An instance of video data can be broken up into segments that can each be analyzed to determine faces and bodies represented therein. The bodies can be analyzed across frames of the segment to determine body tracklets that are consistent across the segment. Associations of faces and bodies can be determined based using relative distances and/or spatial relationships. A subsequent clustering of these associations is performed to attempt to determine consistent associations that correspond to unique individuals. Unique identifiers are determined for each person represented in one or more segments of an instance of video data. Such an approach enables individual representations to be correlated across multiple instances.

Claims (64)

1. A computer-implemented method, comprising:

dividing video data into a plurality of segments, the video data including multiple scenes of a single video feed or multiple video feeds;

analyzing the segments to identify faces and bodies represented in the video data for the segments;

analyzing the bodies represented over a length of individuals of the segments to generate one or more body tracklets for the respective segment;

determining, based at least in part upon spatial relationships between individuals of the faces and individuals of the body tracklets in individual segments, potential associations of the faces and the body tracklets in the respective segment that satisfy a minimum association confidence;

determining, using a clustering algorithm, at least a subset of the potential associations that satisfy a consistency criterion between a first segment and at least a second of the segments; and

outputting unique identifiers for persons identified in the plurality of segments corresponding to the associations that satisfy the consistency criterion.

2. The computer-implemented method of claim 1 , further comprising:

generating a respective feature vector for each association of a face and body tracklet determined in a respective segment; and

processing the feature vectors using the clustering algorithm to generate the subset of associations.

3. The computer-implemented method of claim 1 , further comprising:

analyzing the associations using the consistency criterion, the consistency criterion including at least one of a minimum amount of association, a minimum strength of association, or an incompatible association through one or more segments of the video data.

4. The computer-implemented method of claim 1 , further comprising:

processing frames of the video data using a face recognition algorithm to generate a first set of bounding boxes indicating locations of faces recognized in the video data;

processing frames of the video data using a body recognition algorithm to generate a second set of bounding boxes indicating locations of bodies recognized in the video data; and

determining the associations of faces and the body tracklets based at least in part upon spatial relationships of bounding boxes of the first and second sets.

5. The computer-implemented method of claim 1 , further comprising:

storing the unique identifiers to an identifier repository, wherein a respective unique identifier is associated with multiple instances of video data in which the corresponding person is identified.

6. A computer-implemented method, comprising:

identifying a plurality of faces and a plurality of bodies represented in a segment of video data;

determining a confidence score for individual pairs of faces and bodies based at least in part upon spatial relationships between individuals of the faces and individuals of body tracklets in individual segments;

determining, based on the confidence scores, that a face of the plurality of faces and a body of the plurality of bodies correspond to a single person with at least a minimum confidence;

determining that a correspondence of the face to the body satisfies a consistency criterion over a plurality of the segments of the video data based on the spatial relationship; and

outputting an identifier for a person associated with the face and the body detected in the segment.

7. The computer-implemented method of claim 6 , further comprising:

analyzing a second segment of video data for a different scene of the video data or a second video data;

outputting the identifier for the person as detected in the second segment; and

storing an association of the identifier with the segment of video data and the second segment of video data.

8. The computer-implemented method of claim 6 , further comprising:

obtaining a plurality of instances of video data; and

dividing the instances into a set of video segments including the segment of video data.

9. The computer-implemented method of claim 6 , further comprising:

processing frames of the segment using a face recognition algorithm to identify at least the face represented in the frames.

10. The computer-implemented method of claim 6 , further comprising:

processing frames of the segment using a body detection algorithm to identify at least the body represented in the frames.

11. The computer-implemented method of claim 10 , further comprising:

determining, using an optical flow algorithm, a consistent path of motion of the body across at least a subset of frames of the segment to generate a body tracklet for the representation of the body.

12. The computer-implemented method of claim 11 , further comprising:

determining an association of the face to the body tracklet using at least one of distance or spatial relationship in order to determine that the face and the body correspond to the single person.

13. The computer-implemented method of claim 12 , further comprising:

determining that the correspondence of the face to the body satisfies the consistency criterion by performing clustering of the association over multiple frames of the segment.

14. The computer-implemented method of claim 6 , further comprising:

storing an association of the segment with the identifier for the person in an identifier repository, the identifier capable of being associated with one or more other segments of video data in which a representation of the person was determined.

15. The computer-implemented method of claim 14 , wherein the one or more other segments correspond to at least one of other scenes, feeds, files, or streams of video data.

16. A system, comprising:

at least one processor; and

memory including instructions that, when executed by the at least one processor, cause the system to:

identify a plurality of faces and a plurality of bodies represented in a segment of video data;

determine a confidence score for individual pairs of faces and bodies based at least in part upon spatial relationships between individuals of the faces and individuals of body tracklets in individual segments;

determine, based on the confidence scores, that a face of the plurality of faces and a body of the plurality of bodies correspond to a single person with at least a minimum confidence;

determine that a correspondence of the face to the body satisfies a consistency criterion over a plurality of the segments of the video data based on the spatial relationship; and

output an identifier for a person associated with the face and the body detected in the segment.

17. The system of claim 16 , wherein the instructions when executed further cause the system to:

analyze a second segment of video data for a different scene of the video data or a second video data;

output the identifier for the person as detected in the second segment; and

store an association of the identifier with the segment of video data and the second segment of video data.

18. The system of claim 16 , wherein the instructions when executed further cause the system to:

process frames of the segment using a face recognition algorithm to identify at least the face represented in the frames.

19. The system of claim 16 , wherein the instructions when executed further cause the system to:

process frames of the segment using a body detection algorithm to identify at least the body represented in the frames;

determine, using an optical flow algorithm, a consistent path of motion of the body across at least a subset of frames of the segment to generate a body tracklet for the representation of the body; and

determine an association of the face to the body tracklet using at least one of distance or spatial relationship in order to determine that the face and the body correspond to the single person.

20. The system of claim 16 , wherein the instructions when executed further cause the system to:

store an association of the segment with the identifier for the person in an identifier repository, the identifier capable of being associated with one or more other segments of video data in which a representation of the person was determined.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2018
From: MODOLO, DAVIDE; CHEN, HAO; FILIPPI, ENRICA MARIA; GOULD, STEPHEN; LE MEN, CAMILLE CLAIRE; OLGIATI, ANDREA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 047464/0120 →
Cited By (7)
US 12,219,238 US 12,307,811 US 12,332,938 US 12,412,419 US 12,536,792 US 12,652,455 US 12,718,381