IP Library Granted Patent US 11,600,067
Granted Patent B2
US 11,600,067 · App. 17/016,260 · Granted Mar 7, 2023

Action recognition with high-order interaction through spatial-temporal object tracking

Inventors: Farley Lai (Plainsboro, NJ); Asim Kadav (Jersey City, NJ); Jie Chen (Bellevue, WA)
G06V20/41G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,600,067
App. No.
17/016,260
Granted
Mar 7, 2023
Kind
B2
Abstract

Aspects of the present disclosure describe systems, methods, and structures that provide action recognition with high-order interaction with spatio-temporal object tracking. Image and object features are organized into into tracks, which advantageously facilitates many possible learnable embeddings and intra/inter-track interaction(s). Operationally, our systems, method, and structures according to the present disclosure employ an efficient high-order interaction model to learn embeddings and intra/inter object track interaction across the space and time for AR. Each frame is detected by an object detector to locate visual objects. Those objects are linked through time to form object tracks. The object tracks are then organized and combined with the embeddings as the input to our model. The model is trained to generate representative embeddings and discriminative video features through high-order interaction which is formulated as an efficient matrix operation without iterative processing delay.

Claims (14)

1. A method for determining action recognition in frames of a video through spatio-temporal object tracking, the method comprising:

detecting a plurality of visual objects (O[0] . . . O[n]) in a plurality of frames (F[0] . . . F[n]) of the video;

linking visual objects that are the same through time to form a plurality of object tracks (T[0] . . . T[n]), such that track T[0] includes an image element from each of the frames F[0] . . . F[n], track T[1] consists of objects O[1] from each of the frames F[1], F[2], . . . F[n], track T[2] consists of objects O[2] from each of the frames F[1], F[2], . . . F[n], . . . , and track T[n] consists of objects O[n] from each of the frames F[1], F[2], . . . F[n];

organizing and combining the plurality of object tracks with embeddings, said embeddings including time step embeddings, spatial embeddings, and type/class embeddings;

applying the organized and combined object tracks to a neural network model, said model trained to generate representative embeddings and discriminative video features through high-order interaction formulated as a matrix operation without iterative processing delay.

2. The method of claim 1 wherein the neural network model is a transformer.

3. The method of claim 1 wherein the neural network model includes redesigning input token embeddings for relationship modeling employing a transformer encoder for embedding sequence of image features per frame.

4. The method of claim 1 wherein the neural network model includes redesigning input token embeddings for relationship modeling employing a transformer encoder for embedding sequence of top-K object features per frame.

5. The method of claim 1 wherein the neural network model includes redesigning input token embeddings for relationship modeling using a transformer encoder for embedding sequence of image+object features per frame.

6. The method of claim 1 wherein the applying and organizing includes top 15 objects per frame in transformer based interaction modelling unit with position embeddings.

7. The method of claim 6 further comprising with 2 layers of transformer encoder having 2 parallel heads each.

8. The method of claim 1 , wherein the neural network includes the object tracks are then further organized and input to a model that is trained to generate representative embeddings and discriminative video features through high-order interaction which is formulated as an efficient matrix operation without iterative processing delay.

9. The method of claim 1 , wherein the neural network includes a tracking enabled action recognition process for intra-tracklet and inter-tracklet attention.

10. The method of claim 1 , wherein the neural network includes a video representation input to a tracklet transformer which operationally produces a classification.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2023
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 062403/0866 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2020
From: LAI, FARLEY; KADAV, ASIM; CHEN, JIE
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 053728/0333 →
Continuity (2)
Provisional Application 62899341 · Sep 12, 2019
Related Publication 20210081673A1 · Mar 18, 2021