IP Library › Granted Patent US 12,462,610
Granted Patent B2
US 12,462,610 · App. 17/538,512 · Granted Nov 4, 2025

Video understanding neural network systems and methods using the same

Inventors: Zibo Meng (Palo Alto, CA); Ming Chen (Palo Alto, CA); Chiuman Ho (Palo Alto, CA)
Assignee: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
G06V40/23G06T3/4046G06V10/454G06V10/7715G06V10/82G06V20/41G06V20/46G06N3/045G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,610
App. No.
17/538,512
Granted
Nov 4, 2025
Kind
B2
Abstract

The present disclosure relates to a Temporal Information Aggregation (TIA) neural network block to extract underlying multiscale temporal information. By applying TIA, information in different temporal scales may be effectively extracted. The TIA block may be implemented as a block and thus may be inserted into any architectures. The extracted multi-scale temporal information contributes to the final output as a residual.

Claims (41)

1 . A neural network system implemented by one or more electronic devices, comprising a target neural network block configured to be inserted between a first baseline neural network block and a second baseline neural network block of a baseline neural network, the target neural network block including:

at least one pooling unit to:

receive an input feature map outputted by the first baseline neural network block processing an image,

temporally pooling the input feature map into at least one intermediate feature map;

at least one other processing unit to:

temporally process the at least one intermediate feature map to a residual feature map; and

generate an output feature map configured for inputting into the second baseline neural network block by combining the residual feature map with the input feature map;

wherein temporal information extracted by the target neural network block is different from temporal information extracted by the baseline neural network.

2 . The neural network system of claim 1 , wherein the at least one pooling unit is configured to pooling the input feature map to n intermediate feature maps using n pooling windows of different temporal scales, wherein n is a predetermined positive integer greater than 1.

3 . The neural network system of claim 2 , further comprising:

the baseline neural network, configured to process a plurality of images constructed in a time sequence, including a first baseline block,

wherein the first baseline block includes at least one of a temporal convolutional layer to process a first feature map in temporal domain or a spatial convolutional layer to process the first feature map in spatial domain.

4 . The neural network system of claim 3 , wherein a kernel of the at least one temporal convolutional layer includes a temporal component.

5 . The neural network system of claim 2 , wherein to generate the residual feature map, the target neural network block is further configured to temporally rescale each of the n intermediate feature maps to a predetermined size.

6 . The neural network system of claim 5 , wherein to generate the residual feature map, the at least one other processing unit further includes:

a concatenator to concatenate the n rescaled intermediate feature maps to a concatenated feature map; and

a layer to generate the residual feature map by temporally rescaling the concatenated feature map to a size of the output feature map.

7 . The neural network system of claim 6 , wherein to concatenate the n rescaled intermediate feature maps to the concatenated feature map, the target neural network block is configured to concatenate the n rescale intermediate maps along respective channel axis of the n rescale intermediate maps.

8 . The neural network system of claim 7 , wherein to temporally rescale the concatenated feature map, the layer further includes a convolutional layer to the concatenated feature map to shrink a channel size of the concatenated feature map to a channel size of a second feature map.

9 . The neural network system of claim 1 , wherein to combine the residual feature map with the input feature map, the at least one other processing unit includes a summation unit to elementwise summate the residual feature map with the input feature map.

10 . The neural network system of claim 1 , wherein the baseline neural network further includes a second baseline block connected to the target neural network block in series, and the second baseline block is configured to:

receive the output feature map as a third feature map; and

convert the third feature map to a fourth feature map by process at least one of a spatial feature or a temporal feature of the second feature map.

11 . The neural network system of claim 1 , wherein the baseline neural network employs a 3D ResNet network structure.

12 . A method for analyzing a plurality of images constructed in a time sequence using a target neural network block configured to be inserted between a first baseline neural network block and a second baseline neural network block of a baseline neural network, comprising:

receiving an input feature map outputted by the first baseline neural network block processing an image;

temporally pooling the input feature map into at least one intermediate feature map;

temporally process the at least one intermediate feature map to a residual feature map; and

generating an output feature map configured for inputting into the second baseline neural network block by combining the residual feature map with a second feature map;

wherein temporal information extracted by the target neural network block is different from temporal information extracted by the baseline neural network.

13 . The method of claim 12 , wherein the pooling of the input feature map into at least one intermediate feature map includes temporally pooling the input feature maps to n intermediate feature maps using n pooling windows of different temporal scales, wherein n is a predetermined positive integer greater than 1.

14 . The method of claim 13 , further comprising:

converting a first feature map to a second feature map as the input feature map by respectively processing the first feature map in at least one of temporal domain using a temporal convolutional layer or spatial domain using a spatial convolutional layer.

15 . The method of claim 14 , wherein a kernel of the temporal convolutional layer includes a temporal component.

16 . The method of claim 14 , wherein the combining of the temporal residual feature map with the input feature map includes elementwise summating of the residual feature map with the input feature map.

17 . The method of claim 13 , wherein the generating of the residual feature map includes temporally rescaling each of the n intermediate feature maps to a predetermined size.

18 . The method of claim 17 , wherein the generating of the residual feature map further includes:

concatenating the n rescaled intermediate feature maps to a concatenated feature map; and

generating the residual feature map by temporally rescaling the concatenated feature map to a size of the second feature map.

19 . The method of claim 18 , wherein the temporally rescaling of the concatenated feature map includes applying a convolutional layer to the concatenated feature map to adjust a size of the concatenated feature map to the size of the input feature map.

20 . The method of claim 19 , wherein the temporally rescaling of the concatenated feature map includes applying a convolutional layer to the concatenated feature map to shrink a channel size of the concatenated feature map to a channel size of the second feature map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2021
From: MENG, ZIBO; CHEN, MING; HO, CHIUMAN
To: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
Reel/Frame 058355/0759 →
Continuity (3)
Continuation PCTCN2019122358 · Dec 2, 2019
Provisional Application 62855489 · May 31, 2019
Related Publication 20220157059A1 · May 19, 2022
References Cited (22)
US 20150242746A1 · Rao · 2015 [cited by examiner]
US 20160259994A1 · Ravindran et al. · 2016 [cited by applicant]
US 20180253636A1 · Lee et al. · 2018 [cited by applicant]
US 20200058126A1 · Wang · 2020 [cited by examiner]
CN 106778867 · 2017 [cited by applicant]
CN 108021933 · 2018 [cited by applicant]
CN 109002885 · 2018 [cited by applicant]
CN 109253708 · 2019 [cited by applicant]
CN 109726739 · 2019 [cited by applicant]
JP 2018170003 · 2018 [cited by applicant]
WO 2016197046 · 2016 [cited by applicant]
Karpathy et al., NPL (“Large-scale Video Classification with Convolutional Neural Networks” Published 2014 (Total 8 pages), (Year: 2014). [cited by examiner]
Wu et al., NPL (“Evaluating Two-Stream CNN for Video Classification” Published 2015 (pp. 435-442) (Year: 2015). [cited by examiner]
Guo et al., “Exploiting long-term temporal dynamics for video captioning,” arXiv:2202.10828v1, Feb. 22, 2022. [cited by applicant]
Wu et al., “Action Recognition and Localization with Instance FCNN,” Proceedings of the 2018 IEEE International Conference on Real-time Computing and Robotics, Aug. 2018. [cited by applicant]
CNIPA, First Office Action for CN Application No. 201980096734.5, Jul. 26, 2024. [cited by applicant]
WIPO, International Search Report and Written Opinion for PCT/CN2019/122358, Mar. 4, 2020. [cited by applicant]
EPO, Extended European Search Report for EP Application No. 19930523.6, Jun. 24, 2022. [cited by applicant]
Guo, “Exploiting long-term temporal dynamics for video captioning,” World Wide Web, 2019, vol. 22. [cited by applicant]
Wu et al., “Action Recognition and Localization with Instance FCNN,” Proceedings of the 2018 IEEE International Conference on Real-time Computing and Robotics, 2018. [cited by applicant]
Courtney et al., “Learning From Videos With Deep Convolutional LSTM Networks,” arxiv:1904.04817v1, 2019. [cited by applicant]
Ma et al., “TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition,” Signal Processing: Image Communication, 2018. [cited by applicant]