IP Library Granted Patent US 11,961,333
Granted Patent B2
US 11,961,333 · App. 17/466,117 · Granted Apr 16, 2024

Disentangled representations for gait recognition

Inventors: Xiaoming Liu (Okemos, MI); Ziyuan Zhang (East Lansing, MI)
Assignee: Board of Trustees of Michigan State University
G06V40/25G06V10/761G06V10/774G06V10/82G06V40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,961,333
App. No.
17/466,117
Granted
Apr 16, 2024
Kind
B2
Abstract

Gait, the walking pattern of individuals, is one of the important biometrics modalities. Most of the existing gait recognition methods take silhouettes or articulated body models as gait features. These methods suffer from degraded recognition performance when handling confounding variables, such as clothing, carrying and viewing angle. To remedy this issue, this disclosure proposes to explicitly disentangle appearance, canonical and pose features from RGB imagery. A long short-term memory integrates pose features over time as a dynamic gait feature while canonical features are averaged as a static gait feature. Both of them are utilized as classification features.

Claims (40)

1. A computer-implemented method for identifying a person, comprising:

storing a plurality of features sets, such that each feature set corresponds to a different person, where each feature set includes an identifier for a person, canonical features for the person and gait features for the person;

receiving, by an image processor, a set of images for a given person walking over a period of time;

extracting canonical features of the given person from the set of images using a first neural network, where the canonical features describe body shape of the given person;

extracting gait features of the given person from the set of images using the first neural network and a second neural network, where the gait features describe gait of the given person; and

identifying, by the image processor, the given person by comparing the canonical features of the given person and the gait features of the given person to the plurality of feature sets, where the first neural network and the second neural network are implemented by the image processor.

2. The method of claim 1 wherein extracting canonical features includes extracting a canonical feature from each image in the set of images and averaging the canonical features from each image in the set of images.

3. The method of claim 1 wherein canonical features represent at least one of a shoulder width, a waistline, and a torso to leg ratio.

4. The method of claim 1 further comprises extracting canonical features of the given person from the set of images using a convolutional neural network.

5. The method of claim 1 wherein extracting gait features further comprises

extracting, by the first neural network, pose features of the given person from the set of images, where the pose features describe pose of the given person; and

generating the gait features from the pose features using a long short-term memory.

6. The method of claim 1 further comprises identifying the given person by computing a first cosine similarity between the canonical features of the given person and canonical features from a given feature set, computing a second cosine similarity the gait features of the given person to the gait features from the given feature set, and summing the first cosine similarity score and the second cosine similarity score.

7. The method of claim 1 further comprises

capturing images of a scene using an imaging device, where the scene include the given person walking; and

for each image, segmenting the given person from the image to form the set of images.

8. The method of claim 1 further comprises

receiving a first set of training images for a particular person; and

training the first neural network using the first set of training images in accordance with a first loss function, where the first loss function defines an error between a first image and a second image, such that the first image is reconstructed from appearance features, canonical features and pose features extracted from an image from the first set of training images captured at a given time and the second image is reconstructed from appearance features and canonical feature from the image captured at the given time but pose features from an image from the first set of training images captured at a time subsequent to the given time, wherein the appearance features describe clothes worn by the person.

9. The method of claim 8 further comprises

receiving a second set of training images for the particular person, where appearance features for the particular person extracted from the second set of training images differs from the appearance features for the particular person extracted from the first set of training images;

training the first neural network using the first set of training images and the second set of training images in accordance with a second loss function, where the second loss function defines an error mean of the pose features extracted from the first set of training images and a mean of pose features extracted from a second set of training images.

10. The method of claim 9 further comprises training the first neural network using the first set of training images and the second set of training images in accordance with a third loss function, where the third loss function measures consistency of canonical features across images in the first set of training images; measures consistency of canonical features between the first set of training images and the second set of training images; and measures classification of the particular person using canonical features from at least one of the first set of training images and the second set of training images.

11. The method of claim 10 further comprises

extracting pose features of the particular person from the first set of training images;

generating gait features from the pose features using a long short-term memory;

classifying the particular person with a classifier using the gait features; and

training the second neural network in accordance with a fourth loss function, where the fourth loss function quantifies likelihood that output from the classifier correctly identified the particular person.

12. A non-transitory computer readable medium storing a plurality of feature set and a computer program, where each feature set includes an identifier for a person, canonical features for the person and gait features for the person, the computer program, when executed by a processor, perform to:

receive a set of images for a given person walking over a period of time;

extract canonical features of the given person from the set of images using a first neural network, where the canonical features describe body shape of the given person;

extract gait features of the given person from the set of images using the first neural network and a second neural network, where the gait features describe gait of the given person; and

identifying the given person by comparing the canonical features of the given person and the gait features of the given person to the plurality of feature sets.

13. The non-transitory computer readable medium of claim 12 wherein the computer program further perform to extract a canonical feature from each image in the set of images and average the canonical features from each image in the set of images.

14. The non-transitory computer readable medium of claim 12 wherein the computer program further performs to extract pose features of the given person from the set of images, where the pose features describe pose of the given person; and generate the gait features from the pose features using a long short-term memory.

15. The non-transitory computer readable medium of claim 12 wherein the computer program further performs to identify the given person by computing a first cosine similarity between the canonical features of the given person and canonical features from a given feature set, computing a second cosine similarity the gait features of the given person to the gait features from the given feature set, and summing the first cosine similarity score and the second cosine similarity score.

16. The non-transitory computer readable medium of claim 12 wherein the computer program further performs to capture images of a scene using an imaging device, where the scene include the given person walking; and for each image, segment the given person from the image to form the set of images.

17. The non-transitory computer readable medium of claim 12 wherein the computer program further performs to receive a first set of training images for a particular person; and train the first neural network using the first set of training images in accordance with a first loss function, where the first loss function defines an error between a first image and a second image, such that the first image is reconstructed from appearance features, canonical features and pose features extracted from an image from the first set of training images captured at a given time and the second image is reconstructed from appearance features and canonical feature from the image captured at the given time but pose features from an image from the first set of training images captured at a time subsequent to the given time, wherein the appearance features describe clothes worn by the person.

18. The non-transitory computer readable medium of claim 17 wherein the computer program further performs to receive a second set of training images for the particular person, where appearance features for the particular person extracted from the second set of training images differs from the appearance features for the particular person extracted from the first set of training images; and train the first neural network using the first set of training images and the second set of training images in accordance with a second loss function, where the second loss function defines an error mean of the pose features extracted from the first set of training images and a mean of pose features extracted from a second set of training images.

19. The non-transitory computer readable medium of claim 18 wherein the computer program further performs to train the first neural network using the first set of training images and the second set of training images in accordance with a third loss function, where the third loss function measures consistency of canonical features across images in the first set of training images; measures consistency of canonical features between the first set of training images and the second set of training images; and measure classification of the particular person using canonical features from at least one of the first set of training images and the second set of training images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: LIU, XIAOMING; ZHANG, ZIYUAN
To: BOARD OF TRUSTEES OF MICHIGAN STATE UNIVERSITY
Reel/Frame 066082/0784 →
Continuity (2)
Provisional Application 63074082 · Sep 3, 2020
Related Publication 20220148335A1 · May 12, 2022