IP Library Granted Patent US 12,323,595
Granted Patent B2
US 12,323,595 · App. 17/929,514 · Granted Jun 3, 2025

Processing media using neural networks

Inventors: Dan Grois (Beer-Sheva, IL); Alexander Giladi (Denver, CO)
Assignee: Comcast Cable Communications, LLC
H04N19/127H04N19/103H04N19/137H04N19/159H04N19/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,595
App. No.
17/929,514
Granted
Jun 3, 2025
Kind
B2
Abstract

An encoder may determine a plurality of coding units associated with a frame of a media file and a plurality of prediction units associated with the frame of the media file. The encoder may determine, based on the plurality of coding units associated with the frame and the plurality of prediction units associated with the frame, and based on a training of the encoder using one or more neural networks, that a particular region of the frame can be encoded using one or more encoding characteristics that are different than the encoding characteristics of one or more other particular regions of the frame. The encoder may allocate one or more encoding resources to the particular region of the frame based on the one or more encoding characteristics of the particular region of the frame in order to reduce the overall media bitrate.

Claims (34)

1. A method comprising:

accessing a plurality of frames;

partitioning a frame of the plurality of frames into a plurality of blocks;

determining, using one or more neural networks analyzing one or more textural characteristics of content of the frame, that content of a first block has a textural characteristic and that content of a second block does not have the textural characteristic; and

allocating, based on determining that the content of the first block has the textural characteristic and that content of the second block does not have the textural characteristic, lower encoding resources for the first block of the frame and higher encoding resources for a second block of the frame.

2. The method of claim 1 , wherein the textural characteristic comprises a spatial characteristic.

3. The method of claim 1 , wherein allocating the lower encoding resources for the first block of the frame and higher encoding resources for a second block of the frame comprises allocating fewer bits for one or more motion vectors associated with the first block of the frame than to encoding resources for one or more motion vectors associated with the second block of the frame.

4. The method of claim 1 , wherein the first block of the frame is determined automatically using the one or more neural networks analyzing one or more textural characteristics of the frame.

5. The method of claim 1 , further comprising determining that the first block of the frame can be encoded using one or more first encoding characteristics that are different than second encoding characteristics of the second block of the frame.

6. The method of claim 5 , wherein the determining that the first block of the frame can be encoded using one or more first encoding characteristics that are different than second encoding characteristics of the second block of the frame comprises determining that the one or more textural characteristics of the content in the first block of the frame comprises one or more textures that are different from one or more textures in the second block of the frame.

7. The method of claim 1 , wherein allocating the lower encoding resources for the first block of the frame comprises setting an inter-picture prediction residual signal associated with the first block of the frame to zero.

8. A non-transitory computer-readable medium storing instructions that, when executed, cause:

accessing a plurality of frames;

partitioning a frame of the plurality of frames into a plurality of blocks;

determining, using one or more neural networks analyzing one or more textural characteristics of content of the frame, that content of a first block has a textural characteristic and that content of a second block does not have the textural characteristic; and

allocating, based on determining that the content of the first block has the textural characteristic and that content of the second block does not have the textural characteristic, lower encoding resources for the first block of the frame and higher encoding resources for a second block of the frame.

9. The non-transitory computer-readable medium of claim 8 , wherein the textural characteristic comprises a spatial characteristic.

10. The non-transitory computer-readable medium of claim 8 , wherein the instructions, when executed, cause allocating the lower encoding resources for the first block of the frame and higher encoding resources for a second block of the frame by allocating fewer bits for one or more motion vectors associated with the first block of the frame than to encoding resources for one or more motion vectors associated with the second block of the frame.

11. The non-transitory computer-readable medium of claim 8 , wherein the first block of the frame is determined automatically using the one or more neural networks analyzing one or more textural characteristics of the frame.

12. The non-transitory computer-readable medium of claim 8 , wherein the instructions, when executed, further cause determining that the first block of the frame can be encoded using one or more first encoding characteristics that are different than second encoding characteristics of the second block of the frame.

13. The non-transitory computer-readable medium of claim 12 , wherein the instructions, when executed, cause determining that the first block of the frame can be encoded using one or more first encoding characteristics that are different than second encoding characteristics of the second block of the frame by determining that the one or more textural characteristics of the content in the first block of the frame comprises one or more textures that are different from one or more textures in the second block of the frame.

14. The non-transitory computer-readable medium of claim 8 , wherein the instructions, when executed, cause allocating the lower encoding resources for the first block of the frame by setting an inter-picture prediction residual signal associated with the first block of the frame to zero.

15. A device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the device to:

access a plurality of frames;

partition a frame of the plurality of frames into a plurality of blocks;

determine, using one or more neural networks analyzing one or more textural characteristics of content of the frame, that content of a first block has a textural characteristic and that content of a second block does not have the textural characteristic; and

allocate, based on determining that the content of the first block has the textural characteristic and that content of the second block does not have the textural characteristic, lower encoding resources for the first block of the frame and higher encoding resources for a second block of the frame.

16. The device of claim 15 , wherein the instructions, when executed, cause the device to allocate the lower encoding resources for the first block of the frame and higher encoding resources for a second block of the frame by allocating fewer bits for one or more motion vectors associated with the first block of the frame than to encoding resources for one or more motion vectors associated with the second block of the frame.

17. The device of claim 15 , wherein the first block of the frame is determined automatically using the one or more neural networks analyzing one or more textural characteristics of the frame.

18. The device of claim 15 , wherein the instructions, when executed, further cause the device to determine that the first block of the frame can be encoded using one or more first encoding characteristics that are different than second encoding characteristics of the second block of the frame.

19. The device of claim 18 , wherein the instructions, when executed, cause the device to determine that the first block of the frame can be encoded using one or more first encoding characteristics that are different than second encoding characteristics of the second block of the frame by determining that the one or more textural characteristics of the content in the first block of the frame comprises one or more textures that are different from one or more textures in the second block of the frame.

20. The device of claim 15 , wherein the instructions, when executed, cause the device to allocate the lower encoding resources for the first block of the frame by setting an inter-picture prediction residual signal associated with the first block of the frame to zero.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2022
From: GROIS, DAN; GILADI, ALEXANDER
To: COMCAST CABLE COMMUNICATIONS, LLC
Reel/Frame 060983/0543 →
Continuity (4)
Continuation 17249042 · Feb 18, 2021
Continuation 16736649 · Jan 7, 2020
Provisional Application 62789837 · Jan 8, 2019
Related Publication 20230111773A1 · Apr 13, 2023
References Cited (30)
US 7995649B2 · Zuo · 2011 [cited by examiner]
US 8189933B2 · Holcomb · 2012 [cited by examiner]
US 8243813B2 · Hatabu · 2012 [cited by examiner]
US 10528819B1 · Manmatha · 2020 [cited by examiner]
US 10958908B2 · Grois et al. · 2021 [cited by applicant]
US 20050084007A1 · Lightstone · 2005 [cited by examiner]
US 20140241433A1 · Bosse · 2014 [cited by examiner]
US 20150350655A1 · Huang · 2015 [cited by examiner]
US 20160057418A1 · Lei · 2016 [cited by examiner]
US 20170150150A1 · Thirumalai · 2017 [cited by examiner]
US 20170264902A1 · Ye · 2017 [cited by examiner]
US 20170347107A1 · Teng · 2017 [cited by examiner]
US 20180060691A1 · Bernal · 2018 [cited by examiner]
US 20190230367A1 · Zhu · 2019 [cited by examiner]
US 20190261016A1 · Liu · 2019 [cited by examiner]
US 20190333263A1 · Melkote Krishnaprasad · 2019 [cited by examiner]
US 20200045321A1 · Thirumalai · 2020 [cited by examiner]
US 20200092556A1 · Coelho · 2020 [cited by examiner]
US 20200134837A1 · Varadarajan · 2020 [cited by examiner]
US 20200143457A1 · Manmatha · 2020 [cited by examiner]
US 20210168372A1 · Zhao · 2021 [cited by examiner]
EP 1227684A2 · 2002 [cited by examiner]
GB 2558644A · 2018 [cited by examiner]
Chen Yue et al: “An Overview of Core Coding Tools in the AV1 Video Codec”, 2018 Picture Coding Symposium (PCS), IEEE, Jun. 24, 2018 (Jun. 24, 2018), pp. 41-45, XP033398574. [cited by applicant]
Chichen Fu et al: “Texture Segmentation Based Video Compression Using Convolutional Neural Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Feb. 8, 2018 (Feb. 8, 20… [cited by applicant]
Chun-Man Mak et al: “Enhancing compression rate by just-noticeable distortion model for H.264/AVC”, Circuits and Systems, 2009. ISCAS 2009. IEEE International Symposium on, IEEE, Piscataway, NJ, USA, May 24, 2009 (May 2… [cited by applicant]
Di Chen et al: “AV1 Video Coding Using Texture Analysis With Convolutional Neural Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 25, 2018 (Apr. 25, 2018), XP… [cited by applicant]
Marc Bosch et al: “Segmentation-Based Video Compression Using Texture and Motion Models”, IEEE Journal of Selected Topics in Signal Processing, IEEE, US, vol. 5, No. 7, Nov. 1, 2011 (Nov. 1, 2011), pp. 1366-1377, XP0113… [cited by applicant]
US Patent Application filed Jan. 7, 2020, entitled “Processing Media Using Neural Networks”, U.S. Appl. No. 16/736,649. [cited by applicant]
US Patent Application filed Feb. 18, 2021, entitled “Processing Media Using Neural Networks”, U.S. Appl. No. 17/249,042. [cited by applicant]