IP Library Granted Patent US 12,315,292
Granted Patent B2
US 12,315,292 · App. 16/120,128 · Granted May 27, 2025

Identification of individuals in a digital file using media analysis techniques

Inventors: Balan Rama Ayyar (Oakton, VA); Anantha Krishnan Bangalore (Sterling, VA); Jerome Francois Berclaz (Sunnyvale, CA); Reechik Chatterjee (Washington, DC); Nikhil Kumar Gupta (Oak Hill, VA); Ivan Kovtun (Cupertino, CA); Vasudev Parameswaran (Fremont, CA); Timo Pekka Pylvaenaeinen (Menlo Park, CA); Rajendra Jayantilal Shah (Cupertino, CA)
Assignee: Percipient.AI Inc.
G06V40/162G06F16/738G06F16/784G06V10/761G06V20/30G06V20/41G06V20/49G06V40/165G06V40/168G06V40/172G06V40/173
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,292
App. No.
16/120,128
Granted
May 27, 2025
Kind
B2
Abstract

This description describes a system for identifying individuals within a digital file. The system accesses a digital file describing the movement of unidentified individuals and detects a face for an unidentified individual at a plurality of locations in the video. The system divides the digital file into a set of segments and detects a face of an unidentified individual by applying a detection algorithm to each segment. For each detected face, the system applies a recognition algorithm to extract feature vectors representative of the identity of the detected faces which are stored in computer memory. The system applies a recognition algorithm to query the extracted feature vectors for target individuals by matching unidentified individuals to target individuals, determining a confidence level describing the likelihood that the match is correct, and generating a report to be presented to a user of the system.

Claims (54)

1. In a computer system having at least one processor and computer memory, a method for identifying individuals within a video comprising:

accessing, from the computer memory, a video comprising a plurality of digital frames having an original capture resolution and capturing the movement of one or more unidentified individuals over a period of time;

dividing, in the computer system, at least the majority of the plurality of frames into multiple segments, wherein each segment digitally describes a part of a frame of the video;

in the computer system, adjusting pixel resolution of each segment to a detection resolution;

applying, in the computer system, a detection algorithm configured to detect a face of one or more unidentified individuals within the segment;

in at least some segments, generating a detection bounding box at the detection resolution around each of at least a plurality of the detected faces within that segment, each bounding box having a plurality of vertices;

for each of at least a plurality of the detection bounding boxes, generating a recognition bounding box by mapping each vertex of the detection bounding box from their relative locations within a segment to their proportionally equivalent locations in the original frame of the video, whereby the recognition bounding box is proportionately larger than the detection bounding box;

executing, in the computer system, a recognition algorithm configured to extract, for at least the plurality of the detected faces in each of the at least some segments, a first feature vector representative of the detected face in the associated recognition bounding box wherein the first feature vector is configured to be compared to a second feature vector representative of a target individual's face;

executing in the computer system a search for one or more first feature vectors that substantially match the second feature vector by calculating the distance between the respective feature vectors wherein a substantial match is found when the distance is less than a threshold value;

arranging the substantial matches in accordance with the calculated distance between the associated first feature vector and the second feature vector; and

displaying at least some of the substantial matches for review by a user on a display having at least three graphical elements wherein a first graphical element comprises a scrollable results menu, a second graphical element comprises an analysis view panel, and a third graphical element comprises a timeline tracker, wherein the scrollable results menu comprises a plurality of selectable thumbnail images representative of the substantial matches and a caption describing the location and time at which each associated detection was made, the analysis view panel responsive to the selection of a thumbnail image in the scrollable results menu for displaying a digital file representative of the detection of the face represented by the thumbnail image, and wherein the timeline tracker presents a chronographic record of the time during which a relevant portion of the video was recorded.

2. The method of claim 1 , further comprising:

dividing the video into a set of frames, wherein each set of frames corresponds to a range of timestamps from the period of time during which the video was recorded; and

wherein each segment includes a portion of the frame such that the proportion of a face with respect to the segment is larger relative to the proportion of the face with respect to the frame.

3. The method of claim 1 , wherein the step of executing the recognition algorithm further comprises:

identifying, in the computer system, the face within the bounding box based at least in part on one or more colors of pixels representing physical features of the detected face;

for each background pixel within the bounding box, reducing the influence of each background pixel surrounding the face by normalizing the color of each background pixel within the bounding box; and

following normalization, extracting, in the computer system, the feature vector representative of the detected face.

4. The method of claim 1 , wherein the first feature vector is representative of at least one invariant physical feature of an associated detected face within the segment.

5. The method of claim 1 , further comprising the step of assigning a confidence level for a match wherein the confidence level for a match is inversely related to the calculated distance between the associated first feature vector and the second feature vector.

6. The method of claim 1 , further comprising:

detecting in the computer system, within a plurality of frames, an unidentified individual, wherein the detections are based at least in part on the extracted feature vector for the detected face of the unidentified individual in each frame of the plurality;

determining in the computer system, for pairs of consecutive frames, a distance between the extracted feature vectors;

responsive to determining the distance to be within a threshold distance, generating, for pairs of consecutive frames, an updated feature vector representative of the detected face of an unidentified individual by aggregating the feature vectors from the pair of frames; and

clustering in the computer system, across any plurality of frames of the video, representative feature vectors determined to be within a threshold distance.

7. The method of claim 1 , further comprising:

identifying in the computer system, from one or more frames of the video, frames in which an unidentified individual was present with a second individual;

identifying in the computer system, for each combination of unidentified individuals and second individuals, the number of frames in which both individuals are present; and

assigning in the computer system a label to each combination based on the identified number of frames, the label describing a strength of the relationship between the individuals of the combination.

8. A non-transitory computer readable storage medium comprising stored program code executable by at least one processor, the program code when executed causes the processor to:

access, from computer memory, a video comprising a plurality of frames having an original capture resolution that capture the movement of one or more unidentified individuals over a period of time;

divide at least some of the plurality of frames of the video into one or more sets of segments, wherein each segment describes a part of a frame of the video;

adjust, for each segment, pixel resolution of the segment to a smaller detection resolution such that a detection algorithm detects a face of one more unidentified individuals within the segment;

responsive to the detection algorithm detecting a face, generate a detection bounding box around each of at least a plurality of the detected faces within that segment, each bounding box having a plurality of vertices;

generate, for each of at least a plurality of the bounding boxes, a recognition bounding box by mapping each vertex of the detection bounding box from their relative locations within a segment to their proportionally equivalent locations in the original frame of the video, whereby the recognition bounding box is proportionately larger than the detection bounding box;

execute a recognition algorithm configured to extract a first feature vector representative of the detected face in the recognition bounding box wherein the first feature vector is configured for comparison to a second feature vector representative of a target individual's face, wherein the comparison results from a calculation of the distance between the first feature vector and the second feature vector and a substantial match is found when the distance is less than a threshold value; and

display at least some of the substantial matches for review by a user on a display having at least three graphical elements wherein a first graphical element comprises a scrollable results menu, a second graphical element comprises an analysis view panel, and a third graphical element comprises a timeline tracker, wherein the scrollable results menu comprises a plurality of selectable thumbnail images representative of the substantial matches and a caption describing the location and time at which each associated detection was made, the analysis view panel responsive to the selection of a thumbnail image in the scrollable results menu for displaying a digital file representative of the detection of the face represented by the thumbnail image, and wherein the timeline tracker presents a chronographic record of the time during which a relevant portion of the video was recorded.

9. The non-transitory computer readable storage medium of claim 8 , further comprising stored program code that when executed causes the processor to:

distinguish, by the recognition algorithm, the detected face within the bounding box from background pixels based on one or more colors of pixels representing physical features of the detected face;

for each background pixel within the bounding box, reduce the influence of the environment surrounding the detected face by normalizing the color of each background pixel within the bounding box; and

extract the first feature vector.

10. A system comprising:

an input-output interface, communicatively coupled to at least one processor for at least partly directing storage of data in and retrieval of data from computer memory; and

a non-transitory computer readable storage medium comprising stored program code executable by the at least one processor, the program code when executed causing the processor to:

access, from computer memory, a video comprising one or more frames of pixels capturing the movement of one or more unidentified individuals over a period of time, the frames having an original capture resolution;

divide at least some frames of the video into one or more sets of segments, wherein each segment describes a part of a frame of the video such that the size proportion of a face in the segment increases relative to the proportion of the face in the frame;

adjust, for each segment, pixel resolution of the segment to a detection resolution such that a detection algorithm detects a face of one or more unidentified individuals within the segment;

responsive to the detection algorithm detecting a face, map, for each detected face, their relative locations within a segment at detection resolution to their proportionally equivalent locations in the original frame of the video;

responsive to the mapping, execute a recognition algorithm configured to extract a first feature vector representative of an associated detected face wherein each such first feature vector is configured for comparison to a second feature vector representative of a target individual's face by calculating the distance between the respective feature vectors and a substantial match is found when the distance is less that a threshold value; and

display for review by a user faces of unidentified individuals arranged according to similarity of the comparison between the first feature vector and the second feature vector, the display comprising at least three graphical elements wherein a first graphical element comprises a scrollable results menu, a second graphical element comprises an analysis view panel, and a third graphical element comprises a timeline tracker, wherein the scrollable results menu comprises a plurality of selectable thumbnail images representative of the unidentified individuals and a caption describing the location and time at which each associated detection was made, the analysis view panel responsive to the selection of a thumbnail image in the scrollable results menu for displaying a digital file representative of the detection of the face represented by the thumbnail image, and wherein the timeline tracker presents a chronographic record of the time during which a relevant portion of the video was recorded.

11. The system of claim 10 , wherein the stored program code further comprises program code that when executed causes the processor to:

generate, by the detection algorithm, a bounding box encompassing the detected face, wherein the bounding box demarcates the detected face from a surrounding background environment recorded by the video, distinguish the face within the bounding box from background pixels based on one or more colors of pixels representing physical features of the face;

for each background pixel of the bounding box, reduce the influence of the environment surrounding the face by normalizing the color of each background pixel within the bounding box; and

extract the first feature vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: AYYAR, BALAN RAMA; BANGALORE, ANANTHA KRISHNAN; BERCLAZ, JEROME FRANCOIS; CHATTERJEE, REECHIK; GUPTA, NIKHIL KUMAR; KOVTUN, IVAN; PARAMESWARAN, VASUDEV; PYLVAENAEINEN, TIMO PEKKA; SHAH, RAJENDRA JAYANTILAL
To: PERCIPIENT.AI INC.
Reel/Frame 047156/0284 →
Continuity (2)
Provisional Application 62553725 · Sep 1, 2017
Related Publication 20190073520A1 · Mar 7, 2019
References Cited (43)
US 7295687B2 · Kee et al. · 2007 [cited by applicant]
US 7689011B2 · Luo · 2010 [cited by examiner]
US 8437504B2 · Sugai · 2013 [cited by examiner]
US 9275269B1 · Li · 2016 [cited by examiner]
US 9373024B2 · Liu · 2016 [cited by examiner]
US 9892324B1 · Pachauri · 2018 [cited by examiner]
US 20060204058A1 · Kim · 2006 [cited by examiner]
US 20070110422A1 · Minato · 2007 [cited by examiner]
US 20080144941A1 · Togashi · 2008 [cited by examiner]
US 20080247611A1 · Aisaka et al. · 2008 [cited by applicant]
US 20100287053A1 · Ganong et al. · 2010 [cited by applicant]
US 20110128288A1 · Petrou et al. · 2011 [cited by applicant]
US 20130050492A1 · Lehning · 2013 [cited by applicant]
US 20140270386A1 · Leihs et al. · 2014 [cited by applicant]
US 20160012623A1 · Breckenridge et al. · 2016 [cited by applicant]
US 20160063335A1 · Wang · 2016 [cited by examiner]
US 20160299920A1 · Feng · 2016 [cited by examiner]
US 20170061249A1 · Estrada et al. · 2017 [cited by applicant]
US 20170220887A1 · Fathi et al. · 2017 [cited by applicant]
US 20180114056A1 · Wang et al. · 2018 [cited by applicant]
US 20180157939A1 · Butt · 2018 [cited by examiner]
US 20180268292A1 · Choi et al. · 2018 [cited by applicant]
US 20190073520A1 · Ayyar et al. · 2019 [cited by applicant]
US 20190325595A1 · Stein et al. · 2019 [cited by applicant]
US 20200118292A1 · Estrada et al. · 2020 [cited by applicant]
US 20210133461A1 · Ren · 2021 [cited by examiner]
US 20210407090A1 · Li et al. · 2021 [cited by applicant]
US 20220036194A1 · Sundaresan et al. · 2022 [cited by applicant]
US 20220067394A1 · Suksi et al. · 2022 [cited by applicant]
US 20230020634A1 · Kalman et al. · 2023 [cited by applicant]
CN 114863411 · 2022 [cited by applicant]
WO 2019035771 · 2019 [cited by applicant]
WO 2021146703 · 2021 [cited by applicant]
Liu, W. et al., “SSD: Single Shot MultiBox Detector,” European Conference on Computer Vision, 2016, pp. 21-37. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/US2018/049264, Nov. 28, 2018, 21 pages. [cited by applicant]
Schroff, F. et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, Extended Abstract, 1 page. [cited by applicant]
Bunyarit et al., “Robust Object Detection on Video Surveillance”, 2011 Eighth International Joint Conference on Computer Science and Software Engineering, 2011, retrieved on [Mar. 18, 2021], Retrieved from the internet … [cited by applicant]
Jekel et al., “Classiying Online Profiles on Tinder Using FaceNet Facial Embeddings”, Department of Mechanical & Aerospace Enginnering—University of Florida, Mar. 12, 2018, retrieved on [Mar. 18, 2021]. Retrieved from t… [cited by applicant]
Illa Sucholutsky and Matthias Schonlau, Dept of Statistics and Actuarial Science; University of Waterloo, Waterloo, Ontario, Canada (May 5, 2020). “Soft-Label Dataset Distillation and Text Dataset Distillation.” ISUCHOL… [cited by applicant]
DaAzhi Luo; Guihua Wen; Danyang Li; Yang Hu; Eryang Huan; “Deep-Learning-Based face detection Using Iterative Bounding-Box Regressiion”., Multimed Tools Appl (2018) 77:24663-24680; https://doi.org/1007/s11042-018-5658-5. [cited by applicant]
Beijing University of Posts and Telecommunications, “License Plate Recognition Method and Device” CN114863411; [Jul. 31, 2023; 7:38 AM. [cited by applicant]
Xiang Zhang; Chao Zhao; Hangzai Luo; Wanqing Zhao; Sheng Zhong; Lei Tang; Jinye Peng; Jianping Fan “; School of Information and Technology Northwest Univ, Shaaanxi, 710127, China;” Xian Mictroelectron Technology Institu… [cited by applicant]
Zhang et al., “Automatic Learning for Object Detection”, Neurocomputing, vol. 484, May 1, 2022; pp. 260-272; Received Apr. 10, 2021, Revised Oct. 9, 2021, accepted Feb. 3, 2022, Available online Feb. 7, 2022. Version of… [cited by applicant]
Cited By (1)
US 12,573,201