IP Library › Granted Patent US 12,406,481
Granted Patent B2
US 12,406,481 · App. 18/054,274 · Granted Sep 2, 2025

Video processing using delta distillation

Inventors: Amirhossein Habibian (Amsterdam, NL); Davide Abati (Amsterdam, NL); Haitam Ben Yahia (Diemen, NL)
Assignee: QUALCOMM INCORPORATED
G06V10/7792G06V10/7715G06V10/82G06V20/41G06V20/46G06V20/48G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,481
App. No.
18/054,274
Granted
Sep 2, 2025
Kind
B2
Abstract

Certain aspects of the present disclosure provide techniques and apparatus for processing video content using an artificial neural network. An example method generally includes receiving a video data stream including at least a first frame and a second frame. First features are extracted from the first frame using a teacher neural network. A difference between the first frame and the second frame is determined. Second features are extracted from at least the difference between the first frame and the second frame using a student neural network. A feature map for the second frame is generated based a summation of the first features and the second features. An inference is generated for at least the second frame of the video data stream based on the generated feature map for the second feature.

Claims (64)

1. A processor-implemented method, comprising:

receiving a video data stream including at least a first frame and a second frame;

extracting first features from the first frame using a teacher neural network;

determining a difference between the first frame and the second frame;

extracting second features from at least the difference between the first frame and the second frame using a student neural network;

generating a feature map for the second frame based on a summation of the first features and the second features; and

generating an inference for at least the second frame of the video data stream based on the generated feature map for the second frame.

2. The method of claim 1 , wherein the first frame comprises a key frame in the video data stream and wherein the second frame comprises a non-key-frame in the video data stream.

3. The method of claim 1 , further comprising:

determining a difference between the second frame and a third frame in the video data stream;

extracting third features from at least the difference between the second frame and the third frame using the student neural network;

generating a feature map for the third frame based on a summation of the second features and the third features; and

generating an inference for the third frame of the video data stream based on the generated feature map for the third frame.

4. The method of claim 1 , wherein the teacher neural network comprises a linear network.

5. The method of claim 1 , wherein the student neural network is configured to decompose weights into a lower rank than a rank of weights in the teacher neural network.

6. The method of claim 1 , wherein the student neural network comprises one or more group convolution layers.

7. The method of claim 1 , wherein the teacher neural network comprises a nonlinear network.

8. The method of claim 1 , wherein the student neural network comprises a network with one or more of a reduced number of channels, a reduced spatial resolution, or reduced quantization relative to the teacher neural network.

9. The method of claim 1 , wherein the second features are further extracted from the first frame in combination with the difference between the first frame and the second frame.

10. The method of claim 1 , wherein the student neural network comprises a neural network trained to minimize a loss function based on a difference between an actual change in a feature map between the first frame and the second frame and a predicted change in the feature map between the first frame and the second frame.

11. The method of claim 10 , wherein the loss function is further based on a cost function defined based on a complexity measure of the student neural network and a categorical distribution over a plurality of candidate models.

12. The method of claim 1 , wherein generating the inference comprises identifying one or more objects in the second frame of the video data stream.

13. The method of claim 1 , wherein generating the inference comprises estimating at least one of a pose or a predicted motion of a subject in the video data stream.

14. The method of claim 1 , wherein generating the inference comprises semantically segmenting the video data stream into a plurality of segments associated with different subjects captured in the video data stream.

15. The method of claim 1 , wherein generating the inference comprises mapping the second frame to a code from a plurality of codes in a latent space, and wherein the method further comprises modifying the second frame based on the code in the latent space to which the second frame is mapped.

16. A processor-implemented method, comprising:

receiving a training data set including a plurality of video samples, each video sample of the plurality of video samples including a plurality of frames;

training a teacher neural network based on the training data set;

training a student neural network based on predicted differences between feature maps for successive frames in each video sample and actual differences between feature maps for the successive frames in each video sample; and

deploying the teacher neural network and the student neural network.

17. The method of claim 16 , wherein:

the teacher neural network and the student neural network are trained to minimize a same task-specific objective function, and

the task-specific objective function comprises a function defined based on a weighted delta distribution loss term associated with a difference between actual and predicted changes in feature maps generated for successive frames in a video sample and a weighted cost term associated with a complexity measure for the student neural network.

18. The method of claim 16 , wherein training the student neural network comprises training the student neural network to minimize a loss between an actual difference between outputs generated for successive frames in a video sample in the training data set and a predicted difference between the outputs generated for the successive frames in the video sample.

19. The method of claim 16 , wherein training the student neural network comprises training the student neural network to minimize a cost function defined based on a complexity measure for the student neural network and a categorical distribution over a plurality of candidate models.

20. The method of claim 16 , wherein the teacher neural network comprises a linear network.

21. The method of claim 16 , wherein the student neural network comprises a network configured to decompose weights into a lower rank than a rank of weights of the teacher neural network.

22. A processing system, comprising:

a memory having executable instructions stored thereon; and

a processor configured to execute the executable instructions in order to cause the processing system to:

receive a video data stream including at least a first frame and a second frame;

extract first features from the first frame using a teacher neural network;

determine a difference between the first frame and the second frame;

extract second features from at least the difference between the first frame and the second frame using a student neural network;

generate a feature map for the second frame based on a summation of the first features and the second features; and

generate an inference for at least the second frame of the video data stream based on the generated feature map for the second frame.

23. The processing system of claim 22 , wherein the processor is further configured to cause the processing system to:

determine a difference between the second frame and a third frame in the video data stream;

extract third features from at least the difference between the second frame and the third frame using the student neural network;

generate a feature map for the third frame based on a summation of the second features and the third features; and

generate an inference for the third frame of the video data stream based on the generated feature map for the third frame.

24. The processing system of claim 22 , wherein the teacher neural network comprises a linear network.

25. The processing system of claim 22 , wherein the teacher neural network comprises a nonlinear network, and the student neural network comprises a network with one or more of a reduced number of channels, a reduced spatial resolution, or reduced quantization relative to the teacher neural network.

26. The processing system of claim 22 , wherein the second features are further extracted from the first frame in combination with the difference between the first frame and the second frame.

27. The processing system of claim 22 , wherein the student neural network comprises a neural network trained to minimize a loss function based on a difference between an actual change in a feature map between the first frame and the second frame and a predicted change in the feature map between the first frame and the second frame.

28. The processing system of claim 27 , wherein the loss function is further based on a cost function defined based on a complexity measure of the student neural network and a categorical distribution over a plurality of candidate models.

29. A processing system, comprising:

a memory having executable instructions stored thereon; and

a processor configured to execute the executable instructions in order to cause the processing system to:

receive a training data set including a plurality of video samples, each video sample of the plurality of video samples including a plurality of frames;

train a teacher neural network based on the training data set;

train a student neural network based on predicted differences between feature maps for successive frames in each video sample and actual differences between feature maps for the successive frames in each video sample; and

deploy the teacher neural network and the student neural network.

30. The processing system of claim 29 , wherein in order to train the student neural network, the processor is configured to cause the processing system to train the student neural network to minimize a cost function defined based on a complexity measure for the student neural network and a categorical distribution over a plurality of candidate models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2022
From: HABIBIAN, AMIRHOSSEIN; ABATI, DAVIDE; BEN YAHIA, HAITAM
To: QUALCOMM INCORPORATED
Reel/Frame 061904/0403 →
Continuity (2)
Provisional Application 63264072 · Nov 15, 2021
Related Publication 20230154169A1 · May 18, 2023
References Cited (58)
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20190287515A1 · Li · 2019 [cited by examiner]
US 20210297687A1 · Kawai · 2021 [cited by examiner]
US 20210319232A1 · Perazzi · 2021 [cited by examiner]
US 20220021870A1 · Jiang · 2022 [cited by examiner]
US 20230298330A1 · Chan · 2023 [cited by examiner]
Bonneel N., et al., “Blind Video Temporal Consistency”, acmgraph, ACM Transactions on Graphics, Association for Computing Machinery, vol. 34, No. 6, 2015, pp. 196:1-196:9. [cited by applicant]
Chai Y., “Patchwork: A Patch-Wise Attention Network for Efficient Object Detection and Segmentation in Video Streams”, In ICCV, 2019, pp. 3415-3424. [cited by applicant]
Chen W., et al., “FasterSeg: Searching for Faster Real-Time Semantic Segmentation”, ICLR, arXiv:1912.10917v2 [cs.CV], Jan. 16, 2020, 14 pages. [cited by applicant]
Chen Y., et al., “Memory Enhanced Global-Local Aggregation for Video Object Detection”, In CVPR, 2020, pp. 10337-10346. [cited by applicant]
Cordts M., et al., “The Cityscapes Dataset for Semantic Urban Scene Understanding”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, arXiv:1604.01685v2 [cs.CV], Apr. 7, 2016, pp. 1-29. [cited by applicant]
Courbariaux M., et al., “Training Deep Neural Networks with Low Precision Multiplications”, ICLR, arXiv:1412.7024v5 [cs.LG], Sep. 23, 2015, 10 pages. [cited by applicant]
Denil M., et al., “Predicting Parameters in Deep Learning,” In Neurips, Jun. 3, 2013, pp. 1-9, XP055277618, Retrieved from the Internet: URL:https://arxiv.org/pdf/1306.0543v2.pdf [retrieved on Jun. 3, 2016]. [cited by applicant]
Gou J., et al., “Knowledge Distillation: A Survey”, International Journal of Computer Vision, arXiv:2006.05525v7 [cs.LG], May 20, 2021, pp. 1-36. [cited by applicant]
Gupta S., et al., “Deep Learning with Limited Numerical Precision,” 32nd International Conference on Machine Learning, Jun. 30, 2015, pp. 1737-1746, XP055502076, arXiv:1502.02551v1 [cs.LG] Feb. 9, 2015, Lille, France ab… [cited by applicant]
Habibian A., et al., “Skip-Convolutions for Efficient Video Processing”, arXiv:2104.11487v1 [cs.CV], Apr. 23, 2021, 11 pages. [cited by applicant]
Hie Y., et al., “Channel Pruning for Accelerating Very Deep Neural Networks”, arXiv:1707.06168v2 [cs.CV], Aug. 21, 2017, 10 pages. [cited by applicant]
Hinton G., et al., “Distilling the Knowledge in a Neural Network”, NIPS 2014, arXiv:1503.02531 [stat.ML] Mar. 9, 2015, pp. 1-9, URL:https:/arxiv.org/pdf/1503.02531. [cited by applicant]
Hong Y., et al., “Deep Dual-Resolution Networks for Real-Time and Accurate Semantic Segmentation of Road Scenes”, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, arXiv:2101.06085v2 [cs.CV], pp. 1-12. [cited by applicant]
Hu P., et al., “Real-Time Semantic Segmentation with Fast Attention”, ICRA, arXiv:2007.03815v2 [cs.CV], Jul. 9, 2020, 7 pages. [cited by applicant]
Ju P., et al., “Temporally Distributed Networks for Fast Video Semantic Segmentation”, CVPR, arXiv:2004.01800v2 [cs.CV], Apr. 7, 2020, 10 pages. [cited by applicant]
Jaderberg M., et al., “Speeding up Convolutional Neural Networks with Low Rank Expansions”, Computer Vision and Pattern Recognition, arXiv:1405.3866v1 [cs.CV] May 15, 2014, pp. 1-12, [Online]. Available: https://arxiv.o… [cited by applicant]
Jang E., et al., “Categorical Reparameterization with Gumbel-Softmax”, ICLR 2017 Conference, Cornell University Library, 201 Olin Library Cornell University, Ithaca, NY 14853, Aug. 5, 2017, XP080729237, pp. 1-13, arXiv:… [cited by applicant]
Krishnamoorthi R., “Quantizing Deep Convolutional Networks for Efficient Inference: A Whitepaper”, arXiv:1806.08342v1 [cS.LG], arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,… [cited by applicant]
Lai W-S., et al., “Learning Blind Video Temporal Consistency”, In Proceedings of the European conference on computer vision (ECCV), arXiv:1808.00449v1 [cs.CV], Aug. 1, 2018, pp. 170-185. [cited by applicant]
Lei C., et al., “Blind Video Temporal Consistency via Deep Video Prior”, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020, pp. 1-11. [cited by applicant]
Li H., et al., “Pruning Filters for Efficient Convnets”, ICLR, arXiv:1608.08710v3 [cs.CV], Mar. 10, 2017, pp. 1-13. [cited by applicant]
Liu M., et al., “Looking Fast and Slow: Memory-Guided Mobile Video Object Detection”, arXiv:1903.10172v1 [cs.CV], Mar. 25, 2019, 10 pages. [cited by applicant]
Liu M., et al., “Mobile Video Object Detection with Temporally-Aware Feature Maps”, In CVPR, arXiv:1711.06368v2 [cs.CV], Mar. 28, 2018, 10 pages. [cited by applicant]
Liu Y., et al., “Efficient Semantic Video Segmentation with Per-Frame Inference”, ECCV, 2020, pp. 1-17. [cited by applicant]
Maddison C.J., et al., “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”, ICLR 2017 Conference, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, arXiv:161… [cited by applicant]
Mao H., et al., “Patchnet—Short-Range Template Matching for Efficient Video Processing”, arXiv:2103.07371v1 [cs.CV], Mar. 10, 2021, pp. 1-10. [cited by applicant]
Moons B., et al., “Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces”, In ICCV, 2021, pp. 12229-12238. [cited by applicant]
Nagel M., et al., “Data-Free Quantization through Weight Equalization and Bias Correction”, ICCV, arXiv:1906.04721v3 [cs.LG], Nov. 25, 2019, 13 pages. [cited by applicant]
Orsic M., et al., “In Defense of Pre-Trained Imagenet Architectures for Real-Time Semantic Segmentation of Road-Driving Images”, In CVPR, arXiv:1903.08469v2 [cs.CV], Apr. 12, 2019, 10 pages. [cited by applicant]
Rebol M., et al., “Frame-to-Frame Consistent Semantic Segmentation”, In Joint Austrian Computer Vision And Robotics Workshop (ACVRW), arXiv:2008.00948v3 [cs.CV], Aug. 27, 2020, 11 pages. [cited by applicant]
Ren S., et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Computer Vision and Pattern Recognition (arXiv:1506.01497v3) Jan. 6, 2016, pp. 1-14. [cited by applicant]
Romera E., et al., “ERFNet: Efficient Residual Factorized ConvNet for Real-Time Semantic Segmentation”, IEEE Transactions on Intelligent Transportation Systems, 2017, pp. 1-10. [cited by applicant]
Romero A., et al., “FITNETS: Hints for Thin Deep Nets”, ICLR, arXiv:1412.6550v4 [cs.LG], Mar. 27, 2015, pp. 1-13. [cited by applicant]
Russakovsky O., et al., “ImageNet Large Scale Visual Recognition Challenge”, International Journal of Computer Vision, arXiv:1409.0575v3, Jan. 30, 2015, pp. 1-43. [cited by applicant]
Shrivastava A., et al., “Training Region-Based Object Detectors with Online Hard Example Mining”, In CVPR, arXiv:1604.03540v1 [cs.CV], Apr. 12, 2016, 9 pages. [cited by applicant]
Sibechi R., et al., “Exploiting Temporality for Semi-Supervised Video Segmentation”, In ICCV Workshops, arXiv:1908.11309v1 [cs.CV], Aug. 29, 2019, 9 pages. [cited by applicant]
Tai C., et al., “Convolutional Neural Networks with Low-Rank Regularization”, In ICLR, arXiv:1511.06067v3 [cs.LG], Feb. 14, 2016, pp. 1-11. [cited by applicant]
Tan M., et al., “EfficientDet: Scalable and Efficient Object Detection”, In CVPR, 2020, pp. 10781-10790. [cited by applicant]
Tao A., et al., “Hierarchical Multi-Scale Attention for Semantic Segmentation”, arXiv preprint arXiv:2005.10821v1 [cs.CV], May 21, 2020, pp. 1-11. [cited by applicant]
Wang J., et al., “Deep High-Resolution Representation Learning for Visual Recognition”, IEEE Transactions on Pattern Analysis and Machine Intelligence, Mar. 2020, pp. 1-23. [cited by applicant]
Wang Y., et al., “LEDNET: A Lightweight Encoder-Decoder Network for Real-Time Semantic Segmentation”, In IEEE International Conference on Image Processing, arXiv:1905.02423v3 [cs.CV], May 13, 2019, 5 pages. [cited by applicant]
Yu C., et al., “BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation”, In ECCV, arXiv:1808.00897v1 [cs.CV], Aug. 2, 2018, pp. 1-17. [cited by applicant]
Yu C., et al., “BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation”, International Journal of Computer Vision, arXiv:2004.02147v1 [cs.CV], Apr. 5, 2020, pp. 1-16. [cited by applicant]
Zhang X., et al., “Accelerating Very Deep Convolutional Networks for Classification and Detection”, arXiv:1505.06798v2 [cs.CV], Nov. 18, 2015, pp. 1-14. [cited by applicant]
Zhao H., et al., “ICNet for Real-Time Semantic Segmentation on High-Resolution Images”, In ECCV, arXiv:1704.08545v2 [cs.CV], Aug. 20, 2018, pp. 1-16. [cited by applicant]
Zhao H., et al., “Pyramid Scene Parsing Network”, In CVPR, arXiv:1612.01105v2 [cs.CV], Apr. 27, 2017, 11 pages. [cited by applicant]
Zhu X., et al., “Deep Feature Flow for Video Recognition”, In CVPR, arXiv:1611.07715v2 [cs.CV], Jun. 5, 2017, 13 pages. [cited by applicant]
Zhu X., et al., “Flow-Guided Feature Aggregation for Video Object Detection”, In ICCV, 2017, pp. 408-417. [cited by applicant]
Zhu X., et al., “Towards High Performance Video Object Detection for Mobiles”, arXiv:1804.05830v1 [cs.CV], Apr. 16, 2018, pp. 1-18. [cited by applicant]
Dai X., et al., “General Instance Distillation for Object Detection”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 30, 2021, 10 Pages, XP081929843, figures 1-3. [cited by applicant]
International Search Report and Written Opinion—PCT/US2022/079679—ISA/EPO—Mar. 7, 2023. [cited by applicant]
Ying G., et al., “Better Guider Predicts Future Better: Difference Guided Generative Adversarial Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 7, 2019, pp. … [cited by applicant]