IP Library › Granted Patent US 12,260,674
Granted Patent B2
US 12,260,674 · App. 17/983,741 · Granted Mar 25, 2025

System and method for attention-aware relation mixer for person search

Inventors: Mustansar Fiaz (Abu Dhabi, AE); Hisham Cholakkal (Abu Dhabi, AE); Sanath Narayan (Abu Dhabi, AE); Rao Muhammad Anwer (Abu Dhabi, AE); Fahad Khan (Abu Dhabi, AE)
Assignee: Mohamed bin Zayed University of Artificial Intelligence
G06V40/173G06V10/82G06V20/59H04N7/181
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,674
App. No.
17/983,741
Filed
Nov 9, 2022
Granted
Mar 25, 2025
Kind
B2
Art Unit
2425
USPC
348/159
Abstract

A video system and method for person search includes video cameras for capturing video images, a display device, and a computer system. The computer system including a deep learning network to determine person images, from among the video images, matching a target query person. The deep learning network having a person detection branch, a person re-identification branch, and an attention-aware relation mixer connected to the person detection branch and to the person re-identification branch. The attention-aware relation mixer including a relation mixer having a spatial and channel mixer that performs spatial attention followed by spatial mixing (tokenized multi-layered perceptron) and channel attention followed by channel mixing (channel multi-layered perceptron), and a joint spatio-channel attention layer that utilizes 3D attention weights to modulate 3D spatio-channel region of interest features and aggregate the features with output of the relation mixer. A display device displays matching person images for the person search.

Claims (67)

1. A video system for person search, comprising:

at least one video camera for capturing video images;

a display device; and

a computer system having processing circuitry and memory,

the processing circuitry configured to:

receive a target query person,

perform machine learning using a deep learning network to determine person images, from among the video images, matching the target query person, the deep learning network having

a person detection branch;

a person re-identification branch; and

an attention-aware relation mixer (ARM) connected to the person detection branch and to the person re-identification branch,

the attention-aware relation mixer (ARM) including:

a relation mixer having spatial and channel mixer that performs spatial attention followed by spatial mixing by emphasizing local spatial regions of a person using a spatial attention before globally mixing the local spatial regions across all spatial regions, channel attention followed by channel mixing, and an input-output skip connection configured to perform feature re-using within the relation mixer, and

a joint spatio-channel attention layer that utilizes 3D attention weights to modulate 3D spatio-channel region of interest features and aggregate the features with output of the relation mixer; and

the display device is configured to display matching person images for the person search,

wherein in the deep learning network the person detection branch has a region of interest alignment (RoIAlign) block for region of interest alignment and a shared convolution (res5) block,

the person re-identification branch having a RoIAlign block and a shared convolution block, and

each said branch is connected to the attention-aware relation mixer (ARM) between the respective RoIAlign block and shared convolution block.

2. The video system of claim 1 , wherein the at least one video camera includes a first and a second video cameras,

wherein the first video camera has a first field of view and the second video camera has a second field of view,

wherein the first field of view is non-overlapping with the second field of view, and

wherein the deep learning network determines that at least one first person image from the first video camera and at least one second person image from the second video camera match the target query person.

3. The video system of claim 1 , wherein the relation mixer further comprises a spatial mixer that performs the spatial mixing and a channel mixer for performing the pointwise feature refinement.

4. The video system of claim 3 , wherein the spatial mixer includes a spatial attention layer.

5. The video system of claim 1 , wherein the memory stores video information in association with the captured video images,

wherein the video information includes time information and camera information.

6. The video system of claim 5 , wherein the at least one camera includes a first and a second video cameras,

wherein the first and the second video cameras are located on a building structure,

wherein the deep learning network determines that at least one first person image from the first video camera and at least one second person image from the second video camera match the target query person, and

wherein the display device is further configured to display matching person images in association with date and time information, camera identification information, and identification information for the target query person.

7. The video system of claim 6 , wherein the display device is further configured to display in real time the location and time information of video cameras that identified the target query person.

8. The video system of claim 5 , wherein the at least one camera includes a first and a second video cameras,

wherein the first and the second video cameras are located in a vehicle cabin and the target query person is a driver of the vehicle,

wherein the deep learning network determines that at least one first person image from the first video camera and at least one second person image from the second video camera match the target query person, and

wherein the display device is further configured to display matching person images in association with behavior information of the driver in order to monitor the driver at different points in time.

9. A non-transitory computer readable storage medium storing a computer program for person search, which when executed by processing circuitry performs a method comprising:

receiving a target query person,

performing machine learning using a deep learning network to determine person images, from among video images captured by at least one video camera, that match the target query person, the deep learning network having

a person detection branch;

a person re-identification branch; and

an attention-aware relation mixer (ARM) connected to the person detection branch and to the person re-identification branch,

the attention-aware relation mixer (ARM) including:

a relation mixer having spatial and channel mixer that performs spatial attention followed by spatial mixing by emphasizing local spatial regions of a person using a spatial attention before globally mixing the local spatial regions across all spatial regions, channel attention followed by channel mixing, and an input-output skip connection configured to perform feature re-using within the relation mixer, and

a joint spatio-channel attention layer that utilizes 3D attention weights to modulate 3D spatio-channel region of interest features and aggregate the features with output of the relation mixer; and

displaying the matching person images for the person search,

wherein in the deep learning network the person detection branch has a region of interest alignment (RoIAlign) block for region of interest alignment and a non-shared convolution (res5) block,

the person re-identification branch having a RoIAlign block and a non-shared convolution block, and

each said branch is connected to the attention-aware relation mixer (ARM) between the respective RoIAlign block and non-shared convolution block.

10. The non-transitory computer readable storage medium of claim 9 , wherein the at least one video camera includes a first and a second video cameras,

wherein the first video camera has a first field of view and the second video camera has a second field of view, and

wherein the first field of view is non-overlapping with the second field of view, and

wherein the method further comprises: determining, by the deep learning network, that at least one first person image from the first video camera and at least one second person image from the second video camera match the target query person.

11. The non-transitory computer readable storage medium of claim 9 , wherein the relation mixer further comprises a spatial mixer that performs the spatial mixing and a channel mixer for performing the pointwise feature refinement.

12. The non-transitory computer readable storage medium of claim 11 , wherein the spatial mixer includes a spatial attention layer.

13. The non-transitory computer readable storage medium of claim 9 , wherein the storage medium stores video information in association with the captured video images,

wherein the video information includes time information and camera information.

14. The non-transitory computer readable storage medium of claim 13 , wherein the at least one camera includes a first and a second video cameras,

wherein the first and the second video cameras are located on a building structure,

the method further comprising;

determining, by the deep learning network, that at least one first person image from the first video camera and at least one second person image from the second video camera match the target query person, and

displaying matching person images in association with date and time information, camera identification information, and identification information for the target query person.

15. The non-transitory computer readable storage medium of claim 14 , further comprising:

displaying in real time the location and time information of video cameras that identified the target query person.

16. The non-transitory computer readable storage medium of claim 13 , wherein the at least one camera includes a first and a second video camera

wherein the first and the second video cameras are located in a vehicle cabin and the target query person is a driver of the vehicle,

wherein the method further comprising;

determining, by the deep learning network, that at least one first person image from the first video camera and at least one second person image from the second video camera match the target query person, and

displaying matching person images in association with behavior information of the driver in order to monitor the driver at different points in time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2022
From: FIAZ, MUSTANSAR; CHOLAKKAL, HISHAM; NARAYAN, SANATH; ANWER, RAO MUHAMMAD; KHAN, FAHAD
To: MOHAMED BIN ZAYED UNIVERSITY OF ARTIFICIAL INTELLIGENCE
Reel/Frame 061708/0028 →
Continuity (1)
Related Publication 20240153308A1 · May 9, 2024
References Cited (14)
US 10713794B1 · He · 2020 [cited by examiner]
US 10818028B2 · Chakraborty · 2020 [cited by examiner]
US 20130343642A1 · Kuo et al. · 2013 [cited by applicant]
US 20180174457A1 · Taylor · 2018 [cited by examiner]
US 20200089977A1 · Lakshmi Narayanan · 2020 [cited by examiner]
US 20200134306A1 · Zhu et al. · 2020 [cited by applicant]
CN 113869233A · 2021 [cited by applicant]
CN 114038052A · 2022 [cited by applicant]
CN 114463604A · 2022 [cited by applicant]
Spatial-Attention Location-Aware Multi-Object Tracking (Jun Han , Weixing Li , Feng Pan , Dongdong Zheng , Qi Gao ,. School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China, Proceedings of the… [cited by examiner]
Dense Convolutional Network and Its Application in Medical Image Analysis, Tao Zhou , XinYu Ye, HuiLing Lu , Xiaomin Zheng, Shi Qiu, and YunCan Liu1, School of Computer Science and Engineering, North Minzu University, Y… [cited by examiner]
Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems 34 (2021), Tolstikhin (Year: 2021). [cited by examiner]
Yao, et al. ; Large-scale person re-identification as retrieval ; IEEE International Conference on Multimedia and Expo (ICME) ; Jul. 10-14, 2017 ; Abstract Only ; 4 Pages. [cited by applicant]
Di Chen et al. “Person Search via A Mask-Guided Two-Stream CNN Model,” In: Proceedings of the European conference on computer vision (ECCV), pp. 734-760 (2018). [cited by applicant]