Context-aware video codec based on machine learning
A method, computer system, and computer program product are provided for video encoding and decoding. A plurality of frames of video data are processed using a machine learning model to identify one or more regions of interest in the plurality of frames. The plurality of frames of video data are encoded by a video encoder such that one or more blocks of the plurality of frames of video data corresponding to the one or more regions of interest are prioritized over other blocks, to thereby produce encoded video data.
1 . A computer-implemented method comprising:
generating a mask based on a plurality of frames of video data, wherein the mask indicates one or more regions of interest;
generating, using a machine learning model, at least one additional mask when a threshold amount of change is identified between frames of the plurality of frames of video data;
processing the plurality of frames of video data using the mask and the at least one additional mask to identify the one or more regions of interest in the plurality of frames, wherein the one or more regions of interest are determined dependent on a network bandwidth; and
encoding the plurality of frames of video data by a video encoder such that one or more blocks of the plurality of frames of video data corresponding to the one or more regions of interest are prioritized over other blocks based on metadata generated by the machine learning model, to thereby produce encoded video data.
2 . The computer-implemented method of claim 1 , further comprising:
providing the encoded video data to a recipient device.
3 . The computer-implemented method of claim 2 , further comprising:
providing data indicating the one or more blocks corresponding to the one or more regions of interest to the recipient device.
4 . The computer-implemented method of claim 1 , wherein the one or more regions of interest comprise one or more foreground regions.
5 . The computer-implemented method of claim 1 , wherein the one or more regions of interest are selected from a group of: an upper body and head of a person appearing in the video data, one or more body parts of the person, a face of the person, a mouth and eyes of the person, and an object being held by the person.
6 . The computer-implemented method of claim 5 , wherein encoding the plurality of frames of the video data comprises prioritizing the mouth and eyes of the person over the face of the person by encoding the mouth and eyes of the person with a greater number of bits than that for the face of the person.
7 . The computer-implemented method of claim 5 , wherein encoding the plurality of frames of the video data comprises prioritizing the face of the person over the upper body and head of the person by encoding the face of the person with a greater number of bits than that for the upper body and head of the person.
8 . The computer-implemented method of claim 1 , wherein the one or more regions of interest comprise a portion of the plurality of frames of video data that has a lower brightness compared to other portions of the plurality of frames of video data.
9 . The computer-implemented method of claim 1 , wherein the one or more regions of interest comprise one or more foreground regions in which motion is detected.
10 . The computer-implemented method of claim 1 , wherein the one or more regions of interest comprise text that is included in any of the plurality of frames of the video data.
11 . The computer-implemented method of claim 1 , wherein the one or more regions of interest that are prioritized are encoded, based on the metadata, more quickly than the other blocks.
12 . The computer-implemented method of claim 1 , wherein the one or more regions of interest comprise text in a slide of a presentation.
13 . The computer-implemented method of claim 1 , wherein each of one or more pixels in the at least one additional mask is labeled to indicate a presence or an absence of the one or more regions of interest.
14 . A system comprising:
one or more computer processors;
one or more computer readable storage media; and
program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising instructions to:
generate a mask based on a plurality of frames of video data, wherein the mask indicates one or more regions of interest;
generate, using a machine learning model, at least one additional mask when a threshold amount of change is identified between frames of the plurality of frames of video data;
process the plurality of frames of video data using the mask and the at least one additional mask to identify the one or more regions of interest in the plurality of frames, wherein the one or more regions of interest are determined dependent on a network bandwidth; and
encode the plurality of frames of video data by a video encoder such that one or more blocks of the plurality of frames of video data corresponding to the one or more regions of interest are prioritized over other blocks based on metadata generated by the machine learning model, to thereby produce encoded video data.
15 . The system of claim 14 , wherein the program instructions further comprise instructions to:
provide the encoded video data to a recipient device.
16 . The system of claim 15 , wherein the program instructions further comprise instructions to:
provide data indicating the one or more blocks corresponding to the one or more regions of interest to the recipient device.
17 . The system of claim 14 , wherein the one or more regions of interest are selected from a group of: an upper body and head of a person appearing in the video data, one or more body parts of the person, a face of the person, a mouth and eyes of the person, and an object being held by the person.
18 . One or more non-transitory computer readable storage media having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform operations including:
generating a mask based on a plurality of frames of video data, wherein the mask indicates one or more regions of interest;
generating, using a machine learning model, at least one additional mask when a threshold amount of change is identified between frames of the plurality of frames of video data;
processing the plurality of frames of video data using the mask and the at least one additional mask to identify the one or more regions of interest in the plurality of frames, wherein the one or more regions of interest are determined dependent on a network bandwidth; and
encoding the plurality of frames of video data by a video encoder such that one or more blocks of the plurality of frames of video data corresponding to the one or more regions of interest are prioritized over other blocks based on metadata generated by the machine learning model, to thereby produce encoded video data.
19 . The one or more non-transitory computer readable storage media of claim 18 , wherein the program instructions further cause the computer to:
provide the encoded video data to a recipient device.
20 . The one or more non-transitory computer readable storage media of claim 19 , wherein the program instructions further cause the computer to:
provide data indicating the one or more blocks corresponding to the one or more regions of interest to the recipient device.