IP Library › Granted Patent US 12,518,522
Granted Patent B2
US 12,518,522 · App. 18/334,840 · Granted Jan 6, 2026

Systems and methods for reducing power consumption of executing learning models in vehicle systems

Inventors: Tomaso Trinci (Florence, IT); Tommaso Bianconcini (Florence, IT); Leonardo Taccari (Florence, IT); Leonardo Sarti (Florence, IT); Francesco Sambo (Florence, IT)
Assignee: Verizon Patent and Licensing Inc.
G06V10/82G06V20/41G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,522
App. No.
18/334,840
Granted
Jan 6, 2026
Kind
B2
Abstract

A device may receive video data that includes a plurality of video frames, and may utilize a scheduling policy to divide the plurality of video frames into a first set of video frames and a second set of video frames. The device may process the first set of video frames, with a first convolutional neural network (CNN) model that includes one or more saliency gates, to generate first predictions and saliency maps, and may generate a trained first CNN model based on the first predictions and the saliency maps. The device may process the second set of video frames and the saliency maps, with a second CNN model that includes a saliency propagation module, to generate second predictions, and may generate a trained second CNN model based on the second predictions. The device may perform actions based on the trained first CNN model and the trained second CNN model.

Claims (61)

1 . A method, comprising:

receiving, by a device, video data that includes a plurality of video frames;

utilizing, by the device, a scheduling policy to divide the plurality of video frames into a first set of video frames and a second set of video frames;

processing, by the device, the first set of video frames, with a first convolutional neural network (CNN) model that includes one or more saliency gates, to generate first predictions and saliency maps;

generating, by the device, a trained first CNN model based on the first predictions and the saliency maps;

processing, by the device, the second set of video frames and the saliency maps, with a second CNN model that includes a saliency propagation module, to generate second predictions;

generating, by the device, a trained second CNN model based on the second predictions,

wherein the trained second CNN model is configured to be utilized more than the trained first CNN model without losing accuracy of predictions; and

performing, by the device, one or more actions based on the trained first CNN model and the trained second CNN model.

2 . The method of claim 1 , wherein utilizing the scheduling policy to divide the plurality of video frames into the first set of video frames and the second set of video frames comprises:

selecting a first quantity of the plurality of video frames as the first set of video frames; and

selecting a second quantity of the plurality of video frames as the second set of video frames,

wherein the second quantity is greater than the first quantity.

3 . The method of claim 1 , wherein a first parameter size of the first CNN model is greater than a second parameter size of the second CNN model.

4 . The method of claim 3 , wherein a first input resolution of the first CNN model is greater than a second input resolution of the second CNN model.

5 . The method of claim 1 , wherein each of the saliency maps identifies salient image regions in a video frame of the first set of video frames.

6 . The method of claim 1 , wherein the one or more saliency gates calculate the saliency maps.

7 . The method of claim 1 , wherein each of the one or more saliency gates is provided after a convolutional block of the first CNN model.

8 . A device, comprising:

one or more processors configured to:

receive video data that includes a plurality of video frames;

select a first quantity of the plurality of video frames as a first set of video frames;

select a second quantity of the plurality of video frames as a second set of video frames,

wherein the second quantity is greater than the first quantity;

process the first set of video frames, with a first convolutional neural network (CNN) model that includes one or more saliency gates, to generate first predictions and saliency maps;

generate a trained first CNN model based on the first predictions and the saliency maps;

process the second set of video frames and the saliency maps, with a second CNN model that includes a saliency propagation module, to generate second predictions;

generate a trained second CNN model based on the second predictions,

wherein the trained second CNN model is configured to be utilized more than the trained first CNN model without losing accuracy of predictions; and

perform one or more actions based on the trained first CNN model and the trained second CNN model.

9 . The device of claim 8 , wherein each of the one or more saliency gates calculates one of the saliency maps based on a hidden representation calculated by a convolutional block of the first CNN model and a last latent representation calculated by the first CNN model.

10 . The device of claim 8 , wherein the saliency propagation module injects spatial priors of the saliency maps into the second CNN model and corrects spatial misalignment due to elapsed time.

11 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to:

modify the first quantity of the plurality of video frames or the second quantity of the plurality of video frames based on the trained first CNN model and the trained second CNN model.

12 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to:

modify a quantity of the one or more saliency gates based on the trained first CNN model and the trained second CNN model.

13 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to one or more of:

process real time video data with the trained first CNN model and the trained second CNN model to generate classifications for the real time video data; or

process real time temporal-based data with the trained first CNN model and the trained second CNN model.

14 . The device of claim 8 , wherein the one or more processors, to perform the one or more actions, are configured to:

implement the trained first CNN model and the trained second CNN model at a traffic location or in a vehicle.

15 . A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

receive video data that includes a plurality of video frames;

utilize a scheduling policy to divide the plurality of video frames into a first set of video frames and a second set of video frames;

process the first set of video frames, with a first convolutional neural network (CNN) model that includes one or more saliency gates, to generate first predictions and saliency maps;

generate a trained first CNN model based on the first predictions and the saliency maps;

process the second set of video frames and the saliency maps, with a second CNN model that includes a saliency propagation module, to generate second predictions,

wherein a first parameter size of the first CNN model is greater than a second parameter size of the second CNN model, and

wherein a first input resolution of the first CNN model is greater than a second input resolution of the second CNN model;

generate a trained second CNN model based on the second predictions,

wherein the trained second CNN model is configured to be utilized more than the trained first CNN model without losing accuracy of predictions; and

perform one or more actions based on the trained first CNN model and the trained second CNN model.

16 . The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions, that cause the device to utilize the scheduling policy to divide the plurality of video frames into the first set of video frames and the second set of video frames, cause the device to:

select a first quantity of the plurality of video frames as the first set of video frames; and

select a second quantity of the plurality of video frames as the second set of video frames, wherein the second quantity is greater than the first quantity.

17 . The non-transitory computer-readable medium of claim 15 , wherein each of the saliency maps identifies salient image regions in a video frame of the first set of video frames, and

wherein the one or more saliency gates calculate the saliency maps.

18 . The non-transitory computer-readable medium of claim 15 , wherein each of the one or more saliency gates is provided after a convolutional block of the first CNN model.

19 . The non-transitory computer-readable medium of claim 15 , wherein each of the one or more saliency gates calculates one of the saliency maps based on a hidden representation calculated by a convolutional block of the first CNN model and a last latent representation calculate by the first CNN model.

20 . The non-transitory computer-readable medium of claim 15 , wherein the saliency propagation module injects spatial priors of the saliency maps into the second CNN model and corrects spatial misalignment due to elapsed time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2023
From: TRINCI, TOMASO; BIANCONCINI, TOMMASO; TACCARI, LEONARDO; SARTI, LEONARDO; SAMBO, FRANCESCO
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 063967/0686 →
Continuity (1)
Related Publication 20240420460A1 · Dec 19, 2024
References Cited (33)
US 10327046B1 · Ni · 2019 [cited by examiner]
US 10922548B1 · Huang · 2021 [cited by examiner]
US 12361971B2 · Su · 2025 [cited by examiner]
US 20230154157A1 · Ehteshami Bejnordi · 2023 [cited by examiner]
US 20230185579A1 · Eranpurwala · 2023 [cited by examiner]
US 20240312252A1 · Qiu · 2024 [cited by examiner]
US 20250078220A1 · Ravindran · 2025 [cited by examiner]
Beyer et al., “FlexiViT: One Model for All Patch Sizes,” arXiv:2212.08013v2, Mar. 23, 2023, 29 Pages. [cited by applicant]
Blalock et al., “What is the State of Neural Network Pruning?” arXiv:2003.03033v1, Mar. 6, 2020, 18 Pages. [cited by applicant]
Chen et al., “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” IEEE Journal of Solid-State Circuits, vol. 52, No. 1, 2016, 12 Pages. [cited by applicant]
Dosovitskly et al., “An Image is Worth 16X16 Words: Transformers for Image Recognition at Scale,” Published as a conference paper at ICLR, Jun. 3, 2021, 22 Pages. [cited by applicant]
Fregin et al., “The DriveU Traffic Light Dataset: Introduction and Comparison with Existing Datasets,” 2018 IEEE International Conference on Robotics and Automation (ICRA), May 21-25, 2018, Brisbane, Australia, 8 Pages. [cited by applicant]
Gou et al., “Knowledge Distillation: A Survey,” International Journal of Computer Vision, vol. 129, No. 6, Jun. 2021, 36 Pages. [cited by applicant]
Han et al., “Dynamic Neural Networks: A Survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 11, Dec. 2, 2021, 20 Pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network,” arXiv:1503.02531v1, Mar. 9, 2015, 9 Pages. [cited by applicant]
Howard et al., “arching for MobileNetV3,” arXiv:1905.02244v5, Nov. 20, 2019, 11 Pages. [cited by applicant]
Jain et al., “Accel: A Corrective Fusion Network for Efficient Semantic Segmentation on Video,” arXiv:1807.06667v4, Jul. 5, 2019, 10 Pages. [cited by applicant]
Jetley et al., “Learn to Pay Attention,” Published as a conference paper at ICLR, Apr. 26, 2018, 14 Pages. [cited by applicant]
Li et al., “Towards Streaming Perception,” arXiv:2005.10420v2, Aug. 25, 2020, 39 Pages. [cited by applicant]
Liang et al., “ANT: Adapt Network Across Time for Efficient Video Processing,” Computer Vision Foundation, IEEE Xplore, 2022, 6 Pages. [cited by applicant]
Molchanov et al., “Importance Estimation for Neural Network Pruning,” arXiv:1906.10771v1, Jun. 25, 2019, 11 Pages. [cited by applicant]
Nilsson et al., “Semantic Video Segmentation by Gated Recurrent Flow Propagation,” arXiv: 1612.08871v2, Oct. 2, 2017, 11 Pages. [cited by applicant]
Ren et al., “SBNet: Sparse Blocks Network for Fast Inference,” Computer Vision Foundation, IEEE Xplore, 2018, 10 Pages. [cited by applicant]
Sabet et al., “Temporal Early Exits for Efficient Video Object Detection,” arXiv:2106.11208v1, Jun. 21, 2021, 11 Pages. [cited by applicant]
Scardapane et al., “Why should we add early exits to neural networks?” Cognitive Computation, arXiv:2004.12814v2, Jun. 23, 2020, 23 Pages. [cited by applicant]
Schlemper et al., “Attention gated networks: Learning to leverage salient regions in medical images,” Medical Image Analysis, vol. 53, 2019, 11 Pages. [cited by applicant]
Shkolnik et al., “Robust Quantization: One Model to Rule Them All,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada, 10 Pages. [cited by applicant]
Tung et al., “Similarity-Preserving Knowledge Distillation,” arXiv:1907.09682v2, Aug. 1, 2019, 10 Pages. [cited by applicant]
Verelst et al., “Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference,” Computer Vision Foundation, IEEE Xplore, 2020, 10 Pages. [cited by applicant]
Wang et al., “HAQ: Hardware-Aware Automated Quantization with Mixed Precision,” arXiv:1811.08886v3, Apr. 6, 2019, 10 Pages. [cited by applicant]
Wang et al., “Not All Images are Worth 16x16Words: Dynamic Transformers for Efficient Image Recognition,” 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Sydney, Australia, 16 Pages. [cited by applicant]
Yang et al., “Resolution Adaptive Networks for Efficient Inference,” Computer Vision Foundation, IEEE Xplore, 2020, 10 Pages. [cited by applicant]
Yu et al., “BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning,” Computer Vision Foundation, IEEE Xplore, 2020, 10 Pages. [cited by applicant]