IP Library Granted Patent US 11,483,577
Granted Patent B2
US 11,483,577 · App. 17/205,618 · Granted Oct 25, 2022

Processing of chroma-subsampled video using convolutional neural networks

Inventor: Robert Gonsalves (Wellesley, MA)
Assignee: Avid Technology, Inc.
H04N19/186G06K9/6256G06N3/0445G06N3/08G06T5/002G06V20/40H04N9/77H04N19/46H04N19/85G06T2207/10016G06T2207/10024G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,483,577
App. No.
17/205,618
Granted
Oct 25, 2022
Kind
B2
Abstract

Efficient processing of chroma-subsampled video is performed using convolutional neural networks (CNNs) in which the luma and chroma channels are processed separately. The luma channel is independently convolved and downsampled and, in parallel, the chroma channels are convolved and then merged with the downsampled luma to generate encoded chroma-subsampled video. Further processing of the encoded video that involves deconvolution and upsampling, splitting into two sets of channels, and further deconvolutions and upsampling is used in CNNs to generate decoded chroma-subsampled video in compression-decompression applications, to remove noise from chroma-subsampled video, or to upsample chroma-subsampled video to RGB 444 video. CNNs with separate luma and chroma processing in which the further processing includes additional convolutions and downsampling may be used for object recognition and semantic search in chroma-subsampled video.

Claims (74)

1. A method of generating encoded chroma-subsampled video, the method comprising:

receiving the chroma-subsampled video; and

using an encoding convolutional neural network (CNN) to:

convolve and downsample a luma channel of the chroma-subsampled video;

perform at least one of a convolution and downsampling of chroma channels of the chroma-subsampled video;

merge the convolved and downsampled luma and chroma channels; and

generate the encoded chroma-subsampled video by convolving and downsampling the merged luma and chroma channels.

2. The method of claim 1 , wherein the encoding of the chroma-subsampled video generates a compressed version of the chroma-subsampled video.

3. The method of claim 1 , wherein the encoding convolutional neural network was trained using a training dataset comprising original chroma-subsampled video as input and comparing the original chroma-subsampled video as expected output with chroma-subsampled video generated by using the encoding CNN to encode the original chroma-subsampled video and a corresponding decoding CNN to decode the encoded chroma-subsampled video.

4. The method of claim 3 , wherein the training includes determining matrix values of kernels used by the encoding and decoding CNNs.

5. The method of claim 1 , further comprising using a decoding CNN to generate video represented in RGB 444 color space from the encoded chroma-subsampled video.

6. The method of claim 5 , wherein the encoding CNN and the decoding CNN were trained using an input data set comprising chroma-subsampled video generated by subsampling original video represented in RGB 444 color space and comparing the original video represented in RGB 444 color space with video represented in RGB 444 color space generated by encoding the original video represented in RGB 444 color space using the encoding CNN and decoding the encoded original video represented in RGB 444 color space using the decoding CNN.

7. The method of claim 1 , wherein the received chroma-sub sampled video is noisy, and further comprising using a decoding CNN to generate denoised chroma-subsampled video from the received noisy video.

8. The method of claim 7 , wherein the encoding CNN and the decoding CNN were trained using an input data set comprising noisy chroma-subsampled video generated by adding noise to an original chroma-subsampled video and comparing the original chroma-subsampled video with denoised chroma-subsampled video generated by sequentially encoding the noisy chroma-subsampled video using the encoding CNN and decoding the encoded chroma-subsampled video using the decoding CNN.

9. The method of claim 1 , wherein the encoding CNN further includes steps that convolve and downsample the encoded chroma-subsampled video to generate an identifier of an object depicted in the received chroma-subsampled video.

10. The method of claim 9 , wherein the encoding CNN is trained using a training dataset comprising:

chroma-subsampled video; and

associated with each frame of the chroma-subsampled video a known identifier of an object depicted in the frame; and

wherein for each frame of the training data set provided as input to the encoding CNN, an output of the encoding CNN is compared with its associated object identifier, and a difference between the identifier output by the CNN and the known identifier is fed back into the encoding CNN.

11. The method of claim 1 , wherein the encoding CNN further includes steps that convolve and downsample the encoded chroma-subsampled video to generate an image embedding for each frame of the received chroma-subsampled video, wherein the embedding of a given frame corresponds to a caption for an object depicted in the given frame.

12. The method of claim 11 , wherein the encoding CNN is trained using a training dataset comprising:

chroma-subsampled video; and

associated with each frame of the chroma-subsampled video a caption for an object depicted in the frame; and

wherein for each frame of the training data set provided as input to the encoding CNN:

an image embedding is generated from the frame by the encoding CNN; and

a text embedding is generated from a caption associated with the frame by a text-encoding CNN; and

a difference between the image embedding and the text embedding is fed back into the encoding CNN and the text CNN.

13. A method of decoding encoded chroma-subsampled video, the method comprising:

receiving the encoded chroma-subsampled video; and

using a decoding convolutional neural network (CNN) to:

split convolved and downsampled merged luma and chroma channels of the encoded chroma-sub sampled video into a first set of channels and a second set of channels;

deconvolve and upsample the first set of channels;

process the second set of channels by performing at least one of a deconvolution and upsampling;

output the deconvolved and upsampled first set of channels as a luma channel of decoded chroma-sub sampled video; and

output the processed second set of channels as chroma channels of the decoded chroma-subsampled video.

14. The method of claim 13 , wherein the encoded chroma-subsampled video comprises chroma-subsampled video that has been compressed by an encoding CNN.

15. The method of claim 13 , wherein the encoding and decoding CNNs were trained using a training dataset comprising original chroma-subsampled video as input and comparing the original chroma-subsampled video as expected output with chroma-subsampled video generated by sequentially using the encoding CNN to encode the original chroma-subsampled video and the decoding CNN to decode the encoded chroma-subsampled video.

16. The method of claim 15 , wherein the training includes determining matrix values of kernels used by the encoding and decoding CNNs.

17. A computer program product comprising:

a non-transitory computer-readable medium with computer-readable instructions encoded thereon, wherein the computer-readable instructions, when processed by a processing device instruct the processing device to perform a method of generating encoded chroma-sub sampled video, the method comprising:

receiving the chroma-subsampled video; and

using an encoding convolutional neural network (CNN) to:

convolve and downsample a luma channel of the chroma-subsampled video;

perform at least one of a convolution and downsampling of chroma channels of the chroma-sub sampled video;

merge the convolved and downsampled luma and chroma channels; and

generate the encoded chroma-sub sampled video by convolving and downsampling the merged luma and chroma channels.

18. A computer program product comprising:

a non-transitory computer-readable medium with computer-readable instructions encoded thereon, wherein the computer-readable instructions, when processed by a processing device instruct the processing device to perform a method of decoding encoded chroma-sub sampled video, the method comprising:

receiving the encoded chroma-subsampled video; and

using a decoding convolutional neural network (CNN) to:

split convolved and downsampled merged luma and chroma channels of the encoded chroma-subsampled video into a first set of channels and a second set of channels;

deconvolve and upsample the first set of channels;

process the second set of channels by performing at least one of a deconvolution and upsampling;

output the deconvolved and upsampled first set of channels as a luma channel of decoded chroma-sub sampled video; and

output the processed second set of channels as chroma channels of the decoded chroma-sub sampled video.

19. A system comprising:

a memory for storing computer-readable instructions; and

a processor connected to the memory, wherein the processor, when executing the computer-readable instructions, causes the system to perform a method of generating encoded chroma-subsampled video, the method comprising:

receiving the chroma-subsampled video; and

using an encoding convolutional neural network (CNN) to:

convolve and downsample a luma channel of the chroma-subsampled video;

perform at least one of a convolution and downsampling of chroma channels of the chroma-sub sampled video;

merge the convolved and downsampled luma and chroma channels; and

generate the encoded chroma-sub sampled video by convolving and downsampling the merged luma and chroma channels.

20. A system comprising:

a memory for storing computer-readable instructions; and

a processor connected to the memory, wherein the processor, when executing the computer-readable instructions, causes the system to perform a method of decoding encoded chroma-subsampled video, the method comprising:

receiving the encoded chroma-subsampled video; and

using a decoding convolutional neural network (CNN) to:

split convolved and downsampled merged luma and chroma channels of the encoded chroma-subsampled video into a first set of channels and a second set of channels;

deconvolve and upsample the first set of channels;

process the second set of channels by performing at least one of a deconvolution and upsampling;

output the deconvolved and upsampled first set of channels as a luma channel of decoded chroma-sub sampled video; and

output the processed second set of channels as chroma channels of the decoded chroma-sub sampled video.

Assignments (4)
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 059107/0976) Recorded Nov 8, 2023
From: JPMORGAN CHASE BANK, N.A.
To: AVID TECHNOLOGY, INC.
Reel/Frame 065523/0159 →
PATENT SECURITY AGREEMENT Recorded Nov 8, 2023
From: AVID TECHNOLOGY, INC.
To: SIXTH STREET LENDING PARTNERS, AS ADMINISTRATIVE AGENT
Reel/Frame 065523/0194 →
SECURITY INTEREST Recorded Feb 25, 2022
From: AVID TECHNOLOGY, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 059107/0976 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2021
From: GONSALVES, ROBERT A
To: AVID TECHNOLOGY, INC.
Reel/Frame 055647/0301 →
Continuity (1)
Related Publication 20220303557A1 · Sep 22, 2022