IP Library Granted Patent US 12,367,654
Granted Patent B2
US 12,367,654 · App. 17/795,178 · Granted Jul 22, 2025

Patch based video coding for machines

Inventors: Jill Boyce (Portland, OR); Palanivel Guruva Reddiar (Chandler, AZ); Praveen Prasad (Chandler, AZ)
Assignee: Intel Corporation
G06V10/25G06T3/40G06V20/40G06V40/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,654
App. No.
17/795,178
Granted
Jul 22, 2025
Kind
B2
Abstract

Devices and techniques related to implementing patch based video coding for machines are discussed. Such patch based video coding includes detecting regions of interest in a frame of video, extracting the detected regions of interest to one or more atlases absent the frame at a resolution not less than the resolution of the regions of interest, and encoding the one or more atlases to a bitstream.

Claims (73)

1. A system, comprising:

a memory to store at least a portion of input video; and

processor circuitry coupled to the memory, the processor circuitry to:

detect a plurality of regions of interest for a machine learning operation in a full frame of the input video;

form one or more atlases comprising the regions of interest at a first resolution, wherein a first region of interest of the plurality of regions of interest is in a first atlas;

generate metadata corresponding to the one or more atlases and indicative of a size and location of each of the regions of interest in the full frame of video;

detect a second region of interest in a subsequent frame of the input video;

resize the first atlas and add the second region of interest to the resized first atlas; and

encode the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams absent a representation of the full frame of video at the first resolution or a resolution higher than the first resolution, wherein the processor circuitry to encode the one or more atlases and the metadata comprises the processor circuitry to encode the resized first atlas.

2. The system of claim 1 , the processor circuitry to:

downscale the full frame of video to a downscaled frame having a second resolution less than the first resolution; and

include the downscaled frame of video at the second resolution in the one or more atlases for encode into the one or more bitstreams.

3. The system of claim 1 , wherein the processor circuitry to encode a first region of interest of the plurality of regions of interest comprises the processor circuitry to perform scalable video encode based on the first region of interest and a corresponding region of the full frame of video at a resolution lower than the first resolution.

4. The system of claim 1 , wherein the metadata comprises, for a first region of interest of the plurality of regions of interest, a top left position of the first region in the full frame and a scaling factor.

5. The system of claim 1 , wherein the processor circuitry to detect a first region of interest of the plurality of regions of interest comprises the processor circuitry to:

perform lookahead analysis to detect corresponding subsequent first regions of interest in a plurality of temporally subsequent frames relative to the full frame; and

size the first region of interest to include the first region of interest and all subsequent first regions of interest.

6. The system of claim 1 , wherein the processor circuitry to detect a first region of interest of the plurality of regions of interest comprises the processor circuitry to:

determine a detected region around an object in the first region of interest; and

expand the detected region of the first region of interest to provide a buffer around the region.

7. The system of claim 1 , wherein a first region of interest of the plurality of regions of interest comprises a representation of a face, the processor circuitry to separate the first region of interest into a first atlas and encrypt a first bitstream corresponding to the first atlas.

8. At least one non-transitory machine-readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to code video for machine learning by:

detecting a plurality of regions of interest for a machine learning operation in a full frame of video;

forming one or more atlases comprising the regions of interest at a first resolution, wherein a first region of interest of the plurality of regions of interest is in a first atlas;

generating metadata corresponding to the one or more atlases and indicative of a size and location of each of the regions of interest in the full frame of video;

detecting a second region of interest in a subsequent frame of the video;

resizing the first atlas and adding the second region of interest to the resized first atlas; and

encoding the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams absent a representation of the full frame of video at the first resolution or a resolution higher than the first resolution, wherein encoding the one or more atlases and the metadata comprises encoding the resized first atlas.

9. The non-transitory machine-readable medium of claim 8 , further comprising instructions that, in response to being executed on the computing device, cause the computing device to code video for machine learning by:

downscaling the full frame of video to a downscaled frame having a second resolution less than the first resolution; and

including the downscaled frame of video at the second resolution in the one or more atlases for encode into the one or more bitstreams, wherein encoding a first region of interest of the plurality of regions of interest comprises scalable video encoding based on the first region of interest and a corresponding region of the full frame of video at a resolution lower than the first resolution.

10. The non-transitory machine-readable medium of claim 8 , wherein detecting a first region of interest of the plurality of regions of interest comprises:

performing lookahead analysis to detect corresponding subsequent first regions of interest in a plurality of temporally subsequent frames relative to the full frame; and

sizing the first region of interest to include the first region of interest and all subsequent first regions of interest.

11. A system, comprising:

a memory to store at least a portion of video; and

processor circuitry coupled to the memory, the processor circuitry to:

detect a plurality of first regions of interest for machine learning operations in a full frame of the video;

form an atlas comprising the first regions of interest at a first resolution;

generate metadata corresponding to the atlas and indicative of a size and location of each of the first regions of interest in the full frame;

detect a second region of interest in a subsequent frame of the video;

resize the atlas and add the second region of interest to form a resized atlas; and

encode the resized atlas and the metadata into one or more bitstreams, the one or more bitstreams absent a representation of the full frame of video at the first resolution or a resolution higher than the first resolution.

12. The system of claim 11 , the processor circuitry to:

downscale the full frame to a downscaled frame having a second resolution less than the first resolution; and

include the downscaled frame in the atlas for encode into the one or more bitstreams.

13. The system of claim 11 , wherein the processor circuitry to encode at least one of the first regions of interest comprises the processor circuitry to perform scalable video encode based on the at least one of the first regions of interest and a corresponding region of the full frame at a resolution lower than the first resolution.

14. The system of claim 11 , the processor circuitry to:

perform lookahead analysis to detect corresponding one or more subsequent first regions of interest in a plurality of temporally subsequent frames relative to the full frame; and

size at least one of the first regions of interest to include all subsequent first regions of interest.

15. The system of claim 11 , wherein the processor circuitry to detect at least one of the first regions of interest comprises the processor circuitry to:

determine a detected region around an object in the at least one of the first regions of interest; and

expand the detected region of the at least one of the first regions of interest to provide a buffer.

16. The system of claim 11 , wherein at least one of the first regions of interest comprises a representation of a face, the processor circuitry to encrypt a first bitstream corresponding to the at least one of the first regions of interest.

17. At least one non-transitory machine-readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to code video for machine learning by:

detecting a plurality of first regions of interest for machine learning operations in a full frame of the video;

forming an atlas comprising the first regions of interest at a first resolution;

generating metadata corresponding to the atlas and indicative of a size and location of each of the first regions of interest in the full frame;

detecting a second region of interest in a subsequent frame of the video;

resizing the atlas and adding the second region of interest to form a resized atlas; and

encoding the resized atlas and the metadata into one or more bitstreams, the one or more bitstreams absent a representation of the full frame of video at the first resolution or a resolution higher than the first resolution.

18. The non-transitory machine-readable medium of claim 17 , further comprising instructions that, in response to being executed on the computing device, cause the computing device to code video for machine learning by:

downscaling the full frame to a downscaled frame having a second resolution less than the first resolution; and

include the downscaled frame in the atlas for encode into the one or more bitstreams.

19. The non-transitory machine-readable medium of claim 17 , wherein encoding at least one of the first regions of interest comprises scalable video encoding based on the at least one of the first regions of interest and a corresponding region of the full frame at a resolution lower than the first resolution.

20. The non-transitory machine-readable medium of claim 17 , further comprising instructions that, in response to being executed on the computing device, cause the computing device to code video for machine learning by:

performing lookahead analysis to detect corresponding one or more subsequent first regions of interest in a plurality of temporally subsequent frames relative to the full frame; and

sizing at least one of the first regions of interest to include all subsequent first regions of interest.

21. The non-transitory machine-readable medium of claim 17 , further comprising instructions that, in response to being executed on the computing device, cause the computing device to code video for machine learning by:

determining a detected region around an object in the at least one of the first regions of interest; and

expanding the detected region of the at least one of the first regions of interest to provide a buffer.

22. The non-transitory machine-readable medium of claim 17 , wherein at least one of the first regions of interest comprises a representation of a face, the non-transitory machine-readable medium further comprising instructions that, in response to being executed on the computing device, cause the computing device to code video for machine learning by:

encrypting a first bitstream corresponding to the at least one of the first regions of interest.

Continuity (2)
Provisional Application 63011179 · Apr 16, 2020
Related Publication 20230067541A1 · Mar 2, 2023
References Cited (31)
US 8369633B2 · Lu · 2013 [cited by examiner]
US 10776926B2 · Shrivistava · 2020 [cited by applicant]
US 11025942B2 · Sheikh · 2021 [cited by examiner]
US 20070076957A1 · Wang et al. · 2007 [cited by applicant]
US 20070189623A1 · Ryu · 2007 [cited by applicant]
US 20110200256A1 · Saubat et al. · 2011 [cited by applicant]
US 20120288015A1 · Zhang et al. · 2012 [cited by applicant]
US 20140341280A1 · Yang et al. · 2014 [cited by applicant]
US 20180129934A1 · Tao · 2018 [cited by examiner]
US 20190034734A1 · Yen · 2019 [cited by examiner]
US 20190246130A1 · Sheikh et al. · 2019 [cited by applicant]
US 20200014953A1 · Mammou · 2020 [cited by examiner]
JP 2006033793A · 2006 [cited by applicant]
JP 2014060512A · 2014 [cited by applicant]
KR 1020180135898 · 2018 [cited by applicant]
WO 2014208575A1 · 2014 [cited by applicant]
WO 2018043143A1 · 2018 [cited by applicant]
WO 2018130491A1 · 2018 [cited by applicant]
Unterweger, Andreas, et al. “Building a post-compression region-of-interest encryption framework for existing video surveillance systems: Challenges, obstacles and practical concerns.” Multimedia Systems 22 (2016): 617-… [cited by examiner]
Unterweger, Andreas, et al. “Building a post-compression region-of-interest encryption framework for exisiting video surveillance systems: Challenges, obstacles and practical concerns.” Multimedia Systems 22 (2016): 617… [cited by examiner]
International Preliminary Report on Patentability for PCT Application No. PCT/US2021/027541, dated Oct. 27, 2022. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US2021/027541, dated Jul. 28, 2021. [cited by applicant]
Jung, J. “Immersive Video Activities in MPEG-I: Current Status and Upcoming Challenges”, Orange Labs, Workshop: Computational Imaging with Novel Image Modalities, May 27-28, 2019, Inria, Rennes, France. [cited by applicant]
L. Cui, R. Mekuria, M. Preda and E. S. Jang, “Point-Cloud Compression: Moving Picture Experts Group's New Standard in 2020,” in [cited by applicant]
Extended European Search Report from European Patent Application No. 21788552.4 notified Mar. 5, 2024, 11 pgs. [cited by applicant]
Duan, Ling-Yu, et al., “Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics”, airxiv.org, Cornell University Library, 201 Online Library Cornell University Ithaca, NY, Jan. 13, 2… [cited by applicant]
Salahieh, Basel, et al., “Test Model 4 for Immersive Video”, International Organisation for Standardisation, Brussels, BE, Jan. 2020, 45 pgs. [cited by applicant]
Unterweger, Andreas, et al., “Building a post-compression region-of-interest encryption framework for existing video surveillance systems”, Multimedia Systems 2016, vol. 22, No. 5, 23 pgs. [cited by applicant]
Xia, Sifeng et al., “An Emerging Coding Paradigm VCM: a Scalable Coding Approach Beyond Feature and Signal,” Pecking University, Beijing, China, 2020, 7 pgs. [cited by applicant]
Office Action from Japanese Patent Application No. 2022-553027 notified Jan. 30, 2025, 5 pgs. [cited by applicant]
Notice of Allowance from Japanese Patent Application No. 2022-553027 notified May 15, 2025, 4 pgs. [cited by applicant]