IP Library Granted Patent US 12,499,204
Granted Patent B2
US 12,499,204 · App. 18/060,207 · Granted Dec 16, 2025

Compressed video recognition

Inventors: Yen-Kuang Chen (Palo Alto, CA); Shao-Wen Yang (San Jose, CA); Ibrahima J. Ndiour (Portland, OR); Yiting Liao (Sunnyvale, CA); Vallabhajosyula S. Somayazulu (Portland, OR); Omesh Tickoo (Portland, OR); Srenivas Varadarajan (Bangalore, IN)
Assignees: HYUNDAI MOTOR COMPANY; Kia Corporation
G06F21/44G06F9/4881G06F9/5044G06F9/5066G06F9/5072G06F16/535G06F16/538G06F16/54G06F16/951G06F18/21G06F18/211G06F18/213G06F18/2163G06F18/22G06F18/24G06F18/24143G06F21/45G06F21/53G06F21/6254G06F21/64G06K15/1886G06N3/04G06N3/045G06N3/063G06N3/08G06N5/022G06T7/11G06T7/70G06V10/20G06V10/40G06V10/454G06V10/75G06V10/82G06V10/95G06V10/96G06V20/00G06V30/19173G06V30/274G06V40/161G06V40/20H04L9/0643H04L9/3239H04L67/12H04L67/51H04N19/46H04N19/80H04W4/70G06F18/24323G06F2209/503G06F2209/506G06F2221/2117G06T7/20G06T7/223G06T2207/10016G06T2207/20021G06T2207/20024G06T2207/20052G06T2207/20056G06T2207/20064G06T2207/20084G06T2207/20221G06T2207/30242G06V30/194G06V2201/10H04L9/50H04L67/10H04N19/12H04N19/124H04N19/167H04N19/172H04N19/176H04N19/42H04N19/44H04N19/48H04N19/513H04N19/625H04N19/63H04W12/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,204
App. No.
18/060,207
Granted
Dec 16, 2025
Kind
B2
Abstract

In one embodiment, an apparatus comprises a communication interface and a processor. The communication interface is to communicate with a plurality of devices. The processor is to: receive compressed data from a first device, wherein the compressed data is associated with visual data captured by sensor(s); perform a current stage of processing on the compressed data using a current CNN, wherein the current stage of processing corresponds to one of a plurality of processing stages associated with the visual data, and wherein the current CNN corresponds to one of a plurality of CNNs associated with the plurality of processing stages; obtain an output associated with the current stage of processing; determine, based on the output, whether processing associated with the visual data is complete; if the processing is complete, output a result associated with the visual data; if the processing is incomplete, transmit the compressed data to a second device.

Claims (49)

1 . A device, comprising:

interface circuitry; and

processing circuitry to:

receive, via the interface circuitry, compressed video data, wherein the compressed video data comprises motion vectors and prediction residuals, wherein video data is encoded in the compressed video data using the motion vectors and the prediction residuals;

extract a plurality of features from the compressed video data, wherein at least some of the features are based on the motion vectors and at least some of the features are based on the prediction residuals; and

recognize content in the video data using an artificial neural network, wherein the artificial neural network is trained to recognize the content based on the plurality of features, wherein the artificial neural network comprises a plurality of convolutional neural networks (CNNs), wherein one or more of the CNN are trained to recognize the content in the video data based on at least a subset of the features, and wherein one or more of the CNNs are trained to recognize the content in the video data based on pixel values in the video data,

wherein the plurality of CNNs are configured as a cascaded neural network having multiple decision stages,

wherein a first decision stage uses motion vectors to attempt an early decision,

wherein a subsequent decision stage uses prediction residuals if the first decision stage is unsuccessful, and

wherein a final decision stage uses decompressed video data if the subsequent decision stage is unsuccessful.

2 . The device of claim 1 , wherein:

the compressed video data further comprises transform coefficients, quantization parameters, and macroblock coding modes; and

at least some of the features are based on the transform coefficients, the quantization parameters, or the macroblock coding modes.

3 . The device of claim 1 , wherein the processing circuitry to recognize the content in the video data using the artificial neural network is further to:

recognize the content in the video data using at least some of the plurality of CNNs, wherein the CNNs are used successively until the content in the video data is recognized.

4 . The device of claim 1 , wherein:

the motion vectors represent motion that occurs between frames of the video data; and

the prediction residuals represent differences between the frames of the video data.

5 . The device of claim 1 , wherein the processing circuitry to receive, via the interface circuitry, the compressed video data is further to:

receive the compressed video data over a network; or

receive the compressed video data from a camera.

6 . At least one non-transitory computer-readable storage medium having instructions stored thereon, wherein the instructions, when executed on processing circuitry,

cause the processing circuitry to:

receive compressed video data, wherein the compressed video data comprises motion vectors and prediction residuals, wherein video data is encoded in the compressed video data based at least in part on the motion vectors and the prediction residuals; and

recognize content in the video data using an artificial neural network, wherein the artificial neural network is trained to recognize the content based at least in part on the motion vectors and the prediction residuals,

wherein the artificial neural network comprises a plurality of convolutional neural networks (CNNs), wherein one or more of the CNN are trained to recognize the content in the video data based on at least a subset of the features, and wherein one or more of the CNNs are trained to recognize the content in the video data based on pixel values in the video data,

wherein the plurality of CNNs are configured as a cascaded neural network having multiple decision stages,

wherein a first decision stage uses motion vectors to attempt an early decision,

wherein a subsequent decision stage uses prediction residuals if the first decision stage is unsuccessful, and

wherein a final decision stage uses decompressed video data if the subsequent decision stage is unsuccessful.

7 . The storage medium of claim 6 , wherein the instructions that cause the processing circuitry to recognize the content in the video data using the artificial neural network further cause the processing circuitry to:

extract a plurality of features from the compressed video data, wherein at least some of the features are based on the motion vectors and at least some of the features are based on the prediction residuals; and

recognize the content in the video data using the artificial neural network, wherein the artificial neural network is trained to recognize the content based on the plurality of features.

8 . The storage medium of claim 6 , wherein:

the compressed video data further comprises transform coefficients, quantization parameters, and macroblock coding modes; and

the artificial neural network is trained to recognize the content based at least in part on the transform coefficients, the quantization parameters, or the macroblock coding modes.

9 . The storage medium of claim 8 , wherein the transform coefficients comprise discrete cosine transform (DCT) coefficients.

10 . The storage medium of claim 6 , wherein the artificial neural network comprises one or more convolutional neural networks (CNNs), wherein at least one of the CNNs is trained to recognize the content in the video data based at least in part on the motion vectors or the prediction residuals.

11 . The storage medium of claim 10 , wherein the one or more CNNs comprise a plurality of CNNs, wherein at least one of the CNNs is trained to recognize the content in the video data based on pixel values in the video data.

12 . The storage medium of claim 11 , wherein the instructions that cause the processing circuitry to recognize the content in the video data using the artificial neural network further cause the processing circuitry to:

recognize the content in the video data using at least some of the plurality of CNNs, wherein the CNNs are used successively until the content in the video data is recognized.

13 . The storage medium of claim 6 , wherein:

the motion vectors represent motion that occurs between frames of the video data; and

the prediction residuals represent differences between the frames of the video data.

14 . The storage medium of claim 6 , wherein:

the compressed video data is received over a network; or

the compressed video data is received from a camera.

15 . The storage medium of claim 6 , wherein the video data is encoded in the compressed video data based on a motion-compensated video codec.

16 . The storage medium of claim 15 , wherein the motion-compensated video codec is an H.264 video codec.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: INTEL CORPORATION
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION
Reel/Frame 067737/0094 →
Continuity (5)
Continuation 16948861 · Oct 2, 2020
Continuation 16024356 · Jun 29, 2018
Provisional Application 62691464 · Jun 28, 2018
Provisional Application 62611536 · Dec 28, 2017
Related Publication 20230185895A1 · Jun 15, 2023
References Cited (164)
US 6125212A · Kresch et al. · 2000 [cited by applicant]
US 6879266B1 · Dye et al. · 2005 [cited by applicant]
US 6897858B1 · Hashimoto et al. · 2005 [cited by applicant]
US 7587621B2 · Krauthgamer et al. · 2009 [cited by applicant]
US 7774467B1 · Martin et al. · 2010 [cited by applicant]
US 8533166B1 · Sulieman et al. · 2013 [cited by applicant]
US 8588749B1 · Sadhvani et al. · 2013 [cited by applicant]
US 8638978B2 · Alattar et al. · 2014 [cited by applicant]
US 9143776B2 · Zhang et al. · 2015 [cited by applicant]
US 9330426B2 · Davis · 2016 [cited by applicant]
US 9432430B1 · Klenz · 2016 [cited by applicant]
US 9443065B1 · Schneider et al. · 2016 [cited by applicant]
US 9870621B1 · Frueh et al. · 2018 [cited by applicant]
US 10096038B2 · Ramirez et al. · 2018 [cited by applicant]
US 10154274B2 · Lainema et al. · 2018 [cited by applicant]
US 10235880B2 · Seo · 2019 [cited by applicant]
US 10387179B1 · Hildebrant et al. · 2019 [cited by applicant]
US 10528819B1 · Manmatha et al. · 2020 [cited by applicant]
US 10534655B1 · Kinney, Jr. et al. · 2020 [cited by applicant]
US 10559202B2 · Yang et al. · 2020 [cited by applicant]
US 20030046549A1 · Sakata · 2003 [cited by applicant]
US 20040039254A1 · Stivoric et al. · 2004 [cited by applicant]
US 20050113650A1 · Pacione et al. · 2005 [cited by applicant]
US 20050162525A1 · Koshikawa · 2005 [cited by applicant]
US 20050277466A1 · Lock · 2005 [cited by applicant]
US 20060031102A1 · Teller et al. · 2006 [cited by applicant]
US 20070094458A1 · Suwabe · 2007 [cited by applicant]
US 20090222508A1 · Hubbard · 2009 [cited by applicant]
US 20090285492A1 · Ramanujapuram et al. · 2009 [cited by applicant]
US 20090300623A1 · Bansal et al. · 2009 [cited by applicant]
US 20100030740A1 · Higgins et al. · 2010 [cited by applicant]
US 20100050179A1 · Mohindra et al. · 2010 [cited by applicant]
US 20100125563A1 · Nair et al. · 2010 [cited by applicant]
US 20100208698A1 · Lu et al. · 2010 [cited by applicant]
US 20110123127A1 · Mima et al. · 2011 [cited by applicant]
US 20110231899A1 · Pulier · 2011 [cited by applicant]
US 20110246994A1 · Kimbrel et al. · 2011 [cited by applicant]
US 20120087572A1 · Dedeoglu et al. · 2012 [cited by applicant]
US 20120113244A1 · Nielsen et al. · 2012 [cited by applicant]
US 20120159149A1 · Martin et al. · 2012 [cited by applicant]
US 20120170803A1 · Millar et al. · 2012 [cited by applicant]
US 20120173728A1 · Haskins et al. · 2012 [cited by applicant]
US 20120180877A1 · Pallais · 2012 [cited by applicant]
US 20120185528A1 · Jaudon et al. · 2012 [cited by applicant]
US 20120222084A1 · Beaty et al. · 2012 [cited by applicant]
US 20120257056A1 · Otuka · 2012 [cited by applicant]
US 20120278120A1 · Insko et al. · 2012 [cited by applicant]
US 20120284727A1 · Kodialam et al. · 2012 [cited by applicant]
US 20120290725A1 · Podila · 2012 [cited by applicant]
US 20130006469A1 · Green et al. · 2013 [cited by applicant]
US 20130179371A1 · Jain et al. · 2013 [cited by applicant]
US 20130202044A1 · Kitamura et al. · 2013 [cited by applicant]
US 20130269013A1 · Parry et al. · 2013 [cited by applicant]
US 20140098122A1 · Burley et al. · 2014 [cited by applicant]
US 20140137104A1 · Nelson et al. · 2014 [cited by applicant]
US 20140240591A1 · Rajagopalan et al. · 2014 [cited by applicant]
US 20140245298A1 · Zhou et al. · 2014 [cited by applicant]
US 20140283142A1 · Shepherd et al. · 2014 [cited by applicant]
US 20140300739A1 · Mimar · 2014 [cited by applicant]
US 20140351819A1 · Shah et al. · 2014 [cited by applicant]
US 20140368601A1 · deCharms · 2014 [cited by applicant]
US 20140379377A1 · König et al. · 2014 [cited by applicant]
US 20150086115A1 · Danko · 2015 [cited by applicant]
US 20150104103A1 · Candelore · 2015 [cited by applicant]
US 20150116501A1 · McCoy et al. · 2015 [cited by applicant]
US 20150127379A1 · Sorenson · 2015 [cited by applicant]
US 20150135326A1 · Bailey, Jr. · 2015 [cited by applicant]
US 20150139485A1 · Bourdev · 2015 [cited by applicant]
US 20150160884A1 · Scales et al. · 2015 [cited by applicant]
US 20150169746A1 · Hatami-Hanza · 2015 [cited by applicant]
US 20150269481A1 · Annapureddy · 2015 [cited by examiner]
US 20150271517A1 · Pang et al. · 2015 [cited by applicant]
US 20150302219A1 · Lahteenmaki · 2015 [cited by applicant]
US 20150331995A1 · Zhao et al. · 2015 [cited by applicant]
US 20150341633A1 · Richert · 2015 [cited by applicant]
US 20150381942A1 · Brewer et al. · 2015 [cited by applicant]
US 20160033639A1 · Jung et al. · 2016 [cited by applicant]
US 20160042401A1 · Menendez et al. · 2016 [cited by applicant]
US 20160042767A1 · Araya et al. · 2016 [cited by applicant]
US 20160094480A1 · Kulkarni et al. · 2016 [cited by applicant]
US 20160162320A1 · Singh et al. · 2016 [cited by applicant]
US 20160165248A1 · Lainema et al. · 2016 [cited by applicant]
US 20160173364A1 · Pitio et al. · 2016 [cited by applicant]
US 20160179787A1 · DeLeeuw · 2016 [cited by applicant]
US 20160182707A1 · Gabel · 2016 [cited by applicant]
US 20160234071A1 · Nambiar et al. · 2016 [cited by applicant]
US 20160292856A1 · Niemeijer et al. · 2016 [cited by applicant]
US 20160306849A1 · Curino et al. · 2016 [cited by applicant]
US 20160345260A1 · Johnson et al. · 2016 [cited by applicant]
US 20160350146A1 · Udupi et al. · 2016 [cited by applicant]
US 20160350934A1 · Dey · 2016 [cited by examiner]
US 20160358129A1 · Walton et al. · 2016 [cited by applicant]
US 20160358252A1 · Brakenhoff et al. · 2016 [cited by applicant]
US 20170046613A1 · Paluri et al. · 2017 [cited by applicant]
US 20170098086A1 · Hoernecke et al. · 2017 [cited by applicant]
US 20170112391A1 · Stivoric et al. · 2017 [cited by applicant]
US 20170132511A1 · Gong et al. · 2017 [cited by applicant]
US 20170155662A1 · Courbon et al. · 2017 [cited by applicant]
US 20170156594A1 · Stivoric et al. · 2017 [cited by applicant]
US 20170166126A1 · Sypitkowski · 2017 [cited by applicant]
US 20170169227A1 · Rajcan et al. · 2017 [cited by applicant]
US 20170181645A1 · Mahalingam et al. · 2017 [cited by applicant]
US 20170228599A1 · Juan · 2017 [cited by applicant]
US 20170330179A1 · Song et al. · 2017 [cited by applicant]
US 20170337534A1 · Goeringer et al. · 2017 [cited by applicant]
US 20180025175A1 · Kato · 2018 [cited by applicant]
US 20180039745A1 · Chevalier et al. · 2018 [cited by applicant]
US 20180060703A1 · Fineis et al. · 2018 [cited by applicant]
US 20180082296A1 · Brashers · 2018 [cited by applicant]
US 20180183587A1 · Won et al. · 2018 [cited by applicant]
US 20180184236A1 · Faraone et al. · 2018 [cited by applicant]
US 20180188740A1 · May et al. · 2018 [cited by applicant]
US 20180227240A1 · Liu et al. · 2018 [cited by applicant]
US 20180253826A1 · Milanfar et al. · 2018 [cited by applicant]
US 20180300556A1 · Varerkar et al. · 2018 [cited by applicant]
US 20180330610A1 · Wu · 2018 [cited by applicant]
US 20190026914A1 · Hageman et al. · 2019 [cited by applicant]
US 20190033974A1 · Mu · 2019 [cited by examiner]
US 20190034235A1 · Yang et al. · 2019 [cited by applicant]
US 20190034716A1 · Kamarol et al. · 2019 [cited by applicant]
US 20190042867A1 · Chen et al. · 2019 [cited by applicant]
US 20190042870A1 · Chen et al. · 2019 [cited by applicant]
US 20190042900A1 · Smith et al. · 2019 [cited by applicant]
US 20190043201A1 · Strong · 2019 [cited by examiner]
US 20190045207A1 · Chen · 2019 [cited by examiner]
US 20190114487A1 · Vijayanarasimhan et al. · 2019 [cited by applicant]
US 20190123889A1 · Schmidt-Karaca · 2019 [cited by applicant]
US 20190124495A1 · Subramaniam et al. · 2019 [cited by applicant]
US 20190147327A1 · Martin · 2019 [cited by applicant]
US 20190163896A1 · Balaraman et al. · 2019 [cited by applicant]
US 20190163982A1 · Block · 2019 [cited by applicant]
US 20190164322A1 · Kong et al. · 2019 [cited by applicant]
US 20190310891A1 · Baldasaro et al. · 2019 [cited by applicant]
US 20190349426A1 · Smith et al. · 2019 [cited by applicant]
US 20190354410A1 · Baldasaro et al. · 2019 [cited by applicant]
US 20190364492A1 · Azizi et al. · 2019 [cited by applicant]
US 20210366103A1 · Zhang · 2021 [cited by examiner]
US 20220036302A1 · Cella et al. · 2022 [cited by applicant]
CN 102253989A · 2011 [cited by applicant]
CN 106954068A · 2017 [cited by applicant]
EP 0896295A2 · 1999 [cited by applicant]
GB 2516824A · 2015 [cited by applicant]
WO 2017135889A1 · 2017 [cited by applicant]
Zhang et al. ,“Real-time Action Recognition with Enhanced Motion Vectors CNN's” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 2718-2726 (Year: 2016). [cited by examiner]
“The Visual Computing Database: A Platform for Visual Data Processing and Analysis at Internet Scale”, https://pdfs.semanticscholar.org/1f7d/3bdb06ce547139dd1df83c6756fea9d37ac7.pdf, 2015, 17 pages. [cited by applicant]
Chadha Aaron et al.; “Video Classification With CNNs: Using the Codec as a Spatio-Temporal Activity Sensor;” IEEE Transactions on Circuits and Systems for Video Technology, to Appear; accessed at https://arxiv.org/abs/1… [cited by applicant]
Chu, Hong-Min., et al, “Scheduling in Visual Fog Computing: NP-Completeness and Practical Efficient Solutions”, The Thirty-Second AAAI Conferenceon Artificial Intelligence (AAAI-18), Apr. 26, 2018, 9 pages. [cited by applicant]
Jin et al. “A CNN Cascade for Quality Enhancement of Compressed Depth Images”; Dec. 13, 2017 (Year: 2017). [cited by applicant]
Li, Bin, et al; “A multi-branch convolutional neural network for detecting double JPEG compression;” 3rd International Workshop on Digital Crime and Forensics (IWDCF2017); accessed at https://arxiv.org/abs/1710.05477; O… [cited by applicant]
Lu, Yao., et al. “Optasia: A Relational Platform for Efficient Large-Scale Video Analytics”, 2016 ACM Symposium on Cloud Computing (ACMSoCC), Oct. 2016—microsoft.com, pp. 57-70. [cited by applicant]
Pratt, Harry, et al; “FCNN: Fourier Convolutional Neural Networks;” published in “Machine Learning and Knowledge Discovery in Databases,” accessed at ecmlpkdd2017.ijs.si/papers/paperID11.pdf; Jan. 2017; 16 pages. [cited by applicant]
Sabokrou et al. “Deep Cascade: Cascading 3D Deep Neural Networks for Fast Anomaly Detection and Localization in Crowded Scenes” Apr. 2017, IEEE Transactions on Image processing, vol. 26, No. 4 (Year: 2017). [cited by applicant]
Srinivasan, Vignesh et al.; “On the Robustness of Action Recognition Methods in Compressed and Pixel Domain;” Conference: 6th European Workshop on Visual Information Processing (EUVIP); Oct. 2016, 6 pages. [cited by applicant]
Szegedy, Christian, et al.; “Going deeper with convolutions;” accessed at https://arxiv.org/abs/1409.4842; Sep. 17, 2014; 12 pages. [cited by applicant]
Torfason, Robert et al.; “Towards Image Understanding from Deep Compression without Decoding;” 2018 International Conference on Learning Representations, Vancouver, Canada; Feb. 15, 2018; 17 pages. [cited by applicant]
Ulicny, Matej, et al; “On Using CNN with DCT based Image Data;” Proceedings of the Irish Machine Vision and Image Processing Conference (IMVIP 2017), Maynooth, Ireland, Aug. 2017; pp. 44-51. [cited by applicant]
U.S. Appl. No. 11/256,261-B1 Feb. 2022 Bai; Shi. [cited by applicant]
Verma, Vinay, et al; “DCT-domain Deep Convolutional Neural Networks for Multiple JPEG Compression Classification;” accessed at https://arxiv.org/abs/1712.02313; Dec. 6, 2017; 12 pages. [cited by applicant]
Wang, Yunhe, et al.; “CNNpack: packing convolutional neural networks in the frequency domain;” NIPS'16 Proceedings of the 30th International Conference on Neural Information Processing Systems; Barcelona, Spain, Dec. 20… [cited by applicant]
Wu, Chao-Yuan et al.; “Compressed Video Action Recognition;” accessed at https://arxiv.org/abs/1712.00636; Dec. 2, 2017, 14 pages. [cited by applicant]
Office Action issued in Chinese Patent Application No. 2018800634022 dated Mar. 19, 2025, with English translation. [cited by applicant]
An Fengling et al., “Research on Adaptive Tile based Video Image Compression”, China Excellent Doctoral and Master's Thesis Full text Database (Master's) Information Technology Collection, Version I, pp. 138-342, Jul. 1… [cited by applicant]
Office Action issued in corresponding U.S. Appl. No. 18/496,442 dated Aug. 7, 2024. [cited by applicant]
Office Action issued in corresponding U.S. Appl. No. 18/496,442 dated Nov. 22, 2024. [cited by applicant]