Highlight determination using one or more neural networks
Apparatuses, systems, and techniques are presented to select segments of content. In at least one embodiment, one or more neural networks are used to select one or more video segments from one or more video files based at least in part upon an indication of one or more emotions, of one or more viewers of the one or more video segments, during presentation of the one or more video segments.
1 . One or more processors, comprising:
circuitry to:
use one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;
select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;
cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;
use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and
cause the one or more images to be presented to one or more second viewers based on the intensity.
2 . The one or more processors of claim 1 , wherein the one or more images are part of one or more video segments, and the circuitry is further to select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.
3 . The one or more processors of claim 2 , wherein the circuitry is further to select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.
4 . The one or more processors of claim 2 , wherein the one or more first neural networks are further to determine an indication of the one or more emotions expressed by the viewer using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.
5 . The one or more processors of claim 4 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.
6 . The one or more processors of claim 1 , wherein the circuitry is further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.
7 . A system comprising:
one or more processors to use one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;
select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;
cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;
use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and
cause the one or more images to be presented to one or more second viewers based on the intensity.
8 . The system of claim 7 , wherein the one or more images are part of one or more video segments and the one or more processors are further to select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.
9 . The system of claim 8 , wherein the one or more processors are further to select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.
10 . The system of claim 8 , wherein the one or more first neural networks are further to determine an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.
11 . The system of claim 10 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.
12 . The system of claim 7 , wherein the one or more processors are further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.
13 . A method comprising:
predicting, using one or more first neural networks, a classification of one or more emotions expressed by a viewer in response to one or more images;
selecting one more second neural networks based on the classification of the one or more emotions expressed by the viewer;
causing the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;
using the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and
causing the one or more images to be presented to one or more second viewers based on the intensity.
14 . The method of claim 13 , wherein the one or more images are part of one or more video segments, the method further comprising:
selecting the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.
15 . The method of claim 14 , further comprising: selecting the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.
16 . The method of claim 14 , further comprising: determining an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.
17 . The method of claim 16 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.
18 . The method of claim 13 , further comprising:
providing the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.
19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:
predict, using one or more first neural networks, a classification of one or more emotions expressed by a viewer in response to one or more images;
select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;
cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;
use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and
cause the one or more images to be presented to one or more second viewers based on the intensity.
20 . The non-transitory machine-readable medium of claim 19 , wherein the one or more images are part of one or more video segments and wherein the instructions if performed by the one or more processors further cause the one or more processors to:
select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.
21 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed by the one or more processors further cause the one or more processors to:
select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.
22 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed by the one or more processors further cause the one or more processors to:
determine an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.
23 . The non-transitory machine-readable medium of claim 22 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.
24 . The non-transitory machine-readable medium of claim 19 , wherein the one or more processors are further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.
25 . A video selection system, comprising:
one or more processors to use one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;
select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;
cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;
use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and
cause the one or more images to be presented to one or more second viewers based on the intensity; and
memory for storing network parameters for the one or more first neural networks.
26 . The video selection system of claim 25 , wherein the one or more images are part of one or more video segments and the one or more processors are further to select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to one or more video segments.
27 . The video selection system of claim 26 , wherein the one or more processors are further to select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.
28 . The video selection system of claim 26 , wherein the one or more first neural networks are further to determine an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.
29 . The video selection system of claim 28 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.
30 . The video selection system of claim 25 , wherein the one or more processors are further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.
31 . One or more processors, comprising:
circuitry to use to train one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;
select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;
cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;
use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and
cause the one or more images to be presented to one or more second viewers based on the intensity.