Methods, systems, and media for determining perceptual quality indicators of video content items
Techniques for determining perceptual quality indicators of video content items are provided. In some embodiments, a system including one or more processors executes instructions to: receive a video content item comprising a plurality of frames; determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame; determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame; generate a quality level for each frame based on at least the content quality indicator and the video distortion indicator for the frame; and output an indication of quality for the video content item based on the quality level for each frame.
1 . A method comprising:
receiving, by a computing system, a video content item comprising a plurality of frames;
determining, by the computing system and using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame of the plurality of frames of the video content item;
determining, by the computing system and using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;
generating, by the computing system, an indication of quality for the video content item based on at least the video distortion indicator for one or more frames of the plurality of frames and the content quality indicator for one or more frames of the plurality of frames; and
modifying, by the computing system and based on the indication of quality, the video content item or a transmission strategy associated with the video content item by at least:
adjusting a transmission bitrate for the video content item based on the indication of quality;
adjusting a transmission resolution for the video content item based on the indication of quality;
generating a compressed version of the video content item based on the indication of quality;
applying distortion correction to the video content item based on the indication of quality; or
transcoding the video content item from a first format to a second format based on the indication of quality.
2 . The method of claim 1 , further comprising:
determining, by the computing system and using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; and
generating, by the computing system, a quality level for each frame of the plurality of frames of the video content item based on at least the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for the frame, wherein generating the indication of quality is based on the quality level for each frame of the plurality of frames.
3 . The method of claim 2 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.
4 . The method of claim 1 , further comprising:
determining, by the computing system and using the first subnetwork of the deep neural network, predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item; and
outputting, by the computing system, the predicted content labels.
5 . The method of claim 1 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item.
6 . The method of claim 2 , wherein generating the indication of quality comprises averaging the quality level for each frame of the plurality of frames.
7 . The method of claim 1 , further comprising causing, by the computing system, a video recommendation to be presented based on the indication of quality for the video content item.
8 . The method of claim 7 , wherein the video recommendation includes a recommendation to further compress the video content item based on the indication of quality for the video content item.
9 . The method of claim 7 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item.
10 . A system comprising:
a memory that stores instructions; and
one or more processors that execute the instructions to:
receive a video content item comprising a plurality of frames;
determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame of the plurality of frames of the video content item;
determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;
generate an indication of quality for the video content item based on at least the video distortion indicator for one or more frames of the plurality of frames and the content quality indicator for one or more frames of the plurality of frames; and
modify, based on the indication of quality, the video content item or a transmission strategy associated with the video content item by executing the instructions to at least:
adjust a transmission bitrate for the video content item based on the indication of quality;
adjust a transmission resolution for the video content item based on the indication of quality;
generate a compressed version of the video content item based on the indication of quality;
apply distortion correction to the video content item based on the indication of quality; or
transcode the video content item from a first format to a second format based on the indication of quality.
11 . The system of claim 10 , wherein the one or more processors execute the instructions to:
determine, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; and
generate a quality level for each frame of the plurality of frames of the video content item based on at least the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for the frame, wherein generating the indication of quality is based on the quality level for each frame of the plurality of frames.
12 . The system of claim 11 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.
13 . The system of claim 10 , wherein the one or more processors execute the instructions to:
determine, using the first subnetwork of the deep neural network, predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item; and
output the predicted content labels.
14 . The system of claim 10 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item.
15 . The system of claim 11 , wherein, to generate the indication of quality, the one or more processors execute the instructions to average the quality level for each frame of the plurality of frames.
16 . The system of claim 10 , wherein the one or more processors execute the instructions to cause a video recommendation to be presented based on the indication of quality for the video content item.
17 . The system of claim 16 , wherein the video recommendation includes a recommendation to further compress the video content item based on the indication of quality for the video content item.
18 . The system of claim 16 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item.
19 . Non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:
receive a video content item comprising a plurality of frames;
determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame of the plurality of frames of the video content item;
determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;
generate an indication of quality for the video content item based on at least the video distortion indicator for one or more frames of the plurality of frames and the content quality indicator for one or more frames of the plurality of frames; and
modify, based on the indication of quality, the video content item or a transmission strategy associated with the video content item by executing the instructions to at least:
adjust a transmission bitrate for the video content item based on the indication of quality;
adjust a transmission resolution for the video content item based on the indication of quality;
generate a compressed version of the video content item based on the indication of quality;
apply distortion correction to the video content item based on the indication of quality; or
transcode the video content item from a first format to a second format based on the indication of quality.
20 . The non-transitory computer-readable storage media of claim 19 , wherein the instructions, when executed by the one or more processors, cause the one or more processors to:
determine, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; and
generate a quality level for each frame of the plurality of frames of the video content item based on at least the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for the frame, wherein generating the indication of quality is based on the quality level for each frame of the plurality of frames.