IP Library › Granted Patent US 12,651,460
Granted Patent B2
US 12,651,460 · App. 16/871,667 · Granted Jun 9, 2026

Highlight determination using one or more neural networks

Inventors: Siddhant Pardeshi (Pune, IN); Pranit P. Kothari (Pune, IN); Vinayak Vilas Gaikwad (Pune, IN)
Assignee: NVIDIA Corporation
G06V20/47G06N3/08G06V40/19G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,460
App. No.
16/871,667
Granted
Jun 9, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to select segments of content. In at least one embodiment, one or more neural networks are used to select one or more video segments from one or more video files based at least in part upon an indication of one or more emotions, of one or more viewers of the one or more video segments, during presentation of the one or more video segments.

Claims (68)

1 . One or more processors, comprising:

circuitry to:

use one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;

select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;

cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;

use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and

cause the one or more images to be presented to one or more second viewers based on the intensity.

2 . The one or more processors of claim 1 , wherein the one or more images are part of one or more video segments, and the circuitry is further to select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.

3 . The one or more processors of claim 2 , wherein the circuitry is further to select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.

4 . The one or more processors of claim 2 , wherein the one or more first neural networks are further to determine an indication of the one or more emotions expressed by the viewer using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.

5 . The one or more processors of claim 4 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.

6 . The one or more processors of claim 1 , wherein the circuitry is further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.

7 . A system comprising:

one or more processors to use one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;

select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;

cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;

use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and

cause the one or more images to be presented to one or more second viewers based on the intensity.

8 . The system of claim 7 , wherein the one or more images are part of one or more video segments and the one or more processors are further to select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.

9 . The system of claim 8 , wherein the one or more processors are further to select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.

10 . The system of claim 8 , wherein the one or more first neural networks are further to determine an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.

11 . The system of claim 10 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.

12 . The system of claim 7 , wherein the one or more processors are further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.

13 . A method comprising:

predicting, using one or more first neural networks, a classification of one or more emotions expressed by a viewer in response to one or more images;

selecting one more second neural networks based on the classification of the one or more emotions expressed by the viewer;

causing the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;

using the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and

causing the one or more images to be presented to one or more second viewers based on the intensity.

14 . The method of claim 13 , wherein the one or more images are part of one or more video segments, the method further comprising:

selecting the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.

15 . The method of claim 14 , further comprising: selecting the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.

16 . The method of claim 14 , further comprising: determining an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.

17 . The method of claim 16 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.

18 . The method of claim 13 , further comprising:

providing the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.

19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

predict, using one or more first neural networks, a classification of one or more emotions expressed by a viewer in response to one or more images;

select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;

cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;

use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and

cause the one or more images to be presented to one or more second viewers based on the intensity.

20 . The non-transitory machine-readable medium of claim 19 , wherein the one or more images are part of one or more video segments and wherein the instructions if performed by the one or more processors further cause the one or more processors to:

select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to the one or more video segments.

21 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed by the one or more processors further cause the one or more processors to:

select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.

22 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed by the one or more processors further cause the one or more processors to:

determine an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.

23 . The non-transitory machine-readable medium of claim 22 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.

24 . The non-transitory machine-readable medium of claim 19 , wherein the one or more processors are further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.

25 . A video selection system, comprising:

one or more processors to use one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;

select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;

cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;

use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and

cause the one or more images to be presented to one or more second viewers based on the intensity; and

memory for storing network parameters for the one or more first neural networks.

26 . The video selection system of claim 25 , wherein the one or more images are part of one or more video segments and the one or more processors are further to select the one or more video segments based at least in part upon the intensity or a duration of the one or more emotions expressed by the viewer in response to one or more video segments.

27 . The video selection system of claim 26 , wherein the one or more processors are further to select the one or more video segments based at least in part upon an inferred context of the one or more video segments, the inferred context used to select segments representing coherent events.

28 . The video selection system of claim 26 , wherein the one or more first neural networks are further to determine an indication of the one or more emotions using infrared (IR) images of eyes of one or more users captured during presentation of the one or more video segments.

29 . The video selection system of claim 28 , wherein the one or more video segments correspond to virtual reality (VR) content, and wherein the IR images are captured using one or more IR cameras of one or more VR headsets worn by the one or more users.

30 . The video selection system of claim 25 , wherein the one or more processors are further to provide the one or more images to a highlight aggregator for generating a highlight montage including at least a subset of the one or more images.

31 . One or more processors, comprising:

circuitry to use to train one or more first neural networks to predict a classification of one or more emotions expressed by a viewer in response to one or more images;

select one more second neural networks based on the classification of the one or more emotions expressed by the viewer;

cause the one or more first neural networks to provide the one or more emotions expressed by the viewer to the selected one or more second neural networks;

use the one or more second neural networks to detect an intensity associated with the one or more emotions expressed by the viewer; and

cause the one or more images to be presented to one or more second viewers based on the intensity.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded Mar 25, 2026
From: PARDESHI, SIDDHANT; KOTHARI, PRANIT P.; GAIKWAD, VINAYAK VILAS
To: NVIDIA CORPORATION
Reel/Frame 075228/0899 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2020
From: PARDESHI, SIDDHANT; KOTHARI, PRANIT P.; GAIKWAD, VINAYAK VILAS
To: NVIDIA CORPORATION
Reel/Frame 052652/0327 →
Continuity (1)
Related Publication 20210350139A1 · Nov 11, 2021
References Cited (56)
US 10237615B1 · Gudmundsson · 2019 [cited by examiner]
US 10915798B1 · Zhang · 2021 [cited by examiner]
US 11176484B1 · Dorner · 2021 [cited by examiner]
US 11687770B2 · Nesta · 2023 [cited by examiner]
US 20140323817A1 · el Kaliouby · 2014 [cited by examiner]
US 20150067708A1 · Jensen · 2015 [cited by examiner]
US 20150347903A1 · Saxena · 2015 [cited by examiner]
US 20160366203A1 · Blong · 2016 [cited by examiner]
US 20170171614A1 · el Kaliouby et al. · 2017 [cited by applicant]
US 20170293356A1 · Khaderi · 2017 [cited by examiner]
US 20180189581A1 · Turcot · 2018 [cited by examiner]
US 20180294014A1 · Ekambaram · 2018 [cited by examiner]
US 20190094981A1 · Bradski · 2019 [cited by examiner]
US 20190197330A1 · Mahmoud · 2019 [cited by examiner]
US 20190250934A1 · Kim · 2019 [cited by examiner]
US 20190268660A1 · el Kaliouby · 2019 [cited by examiner]
US 20190286229A1 · Chen · 2019 [cited by examiner]
US 20200053425A1 · Gudmundsson et al. · 2020 [cited by applicant]
US 20200134424A1 · Chen · 2020 [cited by examiner]
US 20200170560A1 · Zakariaie · 2020 [cited by examiner]
US 20200285848A1 · Price · 2020 [cited by examiner]
US 20200349429A1 · Vendrow · 2020 [cited by examiner]
US 20200387713A1 · Kanthan · 2020 [cited by examiner]
US 20210125054A1 · Banik · 2021 [cited by examiner]
US 20210306173A1 · Krikunov · 2021 [cited by examiner]
US 20210326574A1 · Wu · 2021 [cited by examiner]
US 20210326586A1 · Sorci · 2021 [cited by examiner]
US 20220121884A1 · Zadeh · 2022 [cited by examiner]
CN 102842327A · 2012 [cited by applicant]
CN 103609128A · 2014 [cited by applicant]
CN 109154861A · 2019 [cited by applicant]
CN 110769314A · 2020 [cited by applicant]
Feffer, M., Rudovic, O.(., Picard, R.W. (2018). A Mixture of Personalized Experts for Human Affect Estimation. In: Perner, P. (eds) Machine Learning and Data Mining in Pattern Recognition. MLDM 2018. Lecture Notes in Co… [cited by examiner]
Zhang et al., “Adaptive 3D facial action intensity estimation and emotion recognition” Expert Systems with Applications vol. 42, Issue 3, Feb. 15, 2015, pp. 1446-1464. (Year: 2015). [cited by examiner]
Brans et al., “Intensity and Duration of Negative Emotions: Comparing the Role of Appraisals and Regulation Strategies,” PLoS One, Mar. 2014, 13 pages. [cited by applicant]
Chorowski et al., Unsupervised Speech Representation Learning using Wavenet Autoencoders, Sep. 11, 2019, 13 pages. [cited by applicant]
Dilokthanakul et al., “Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders,” Nov. 8, 2016, 12 pages. [cited by applicant]
Donahue et al., “DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition,” Jan. 2014, 9 pages. [cited by applicant]
Hickson et al., “Eyemotion: Classifying Facial Expressions in VR using eye-tracking Cameras,” IEEE Winter Conference on Applications of Computer Vision, 2019, 10 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Jacobs et al., “Adaptive Mixtures of Local Experts,” Neural computation , 3 (1), 1991, pp. 79-87. [cited by applicant]
Ng et al., “Addition to the Internet and Online Gaming,” CyberPsychology & Behavior, 8(2): Nov. 2, 2005, 5 pages. [cited by applicant]
Redmon et al., “YOLO9000: Better, Faster, Stronger,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Shazeer et al., “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,” ICLR, Jan. 23, 2017, 19 pages. [cited by applicant]
Simonyan et al., “Two-Stream Convolutional Networks for Action Recognition in Videos,” Advances in Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Sra et al., “Breathvr: Leveraging Breathing as a Directly Controlled Interface for Virtual Reality Games,” Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems CHI, Apr. 21-26, 2018, 12 pages. [cited by applicant]
Szymkowiak, “Google Brain's New Super Fast and Highly Accurate AI: The Mixture of Experts Layer,” retrieved from Internet, https://medium.com/@thoszymkowiak/google-brains-new-super-fast-and-highly-accurate -ai-the-mixtu… [cited by applicant]
Tone et al., “The Attraction of Online Games: An Important Factor for Internet Addiction,” Computers in Human Behavior 30, 2014, 7 pages. [cited by applicant]
Torabi et al., “Action Classification and Highlighting in Videos,” Aug. 31, 2017, 12 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106673.3, mailed Jul. 27, 2023, 3 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106673.3, mailed Dec. 6, 2023, 4 pages. [cited by applicant]
Office Action for Chinese Application No. 202110511245.0, mailed Jun. 26, 2024, 24 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106673.3, mailed Jun. 26, 2024, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2106673.3, mailed Feb. 24, 2023, 4 pages. [cited by applicant]
Office Action for Chinese Application No. 202110511245.0, mailed Dec. 26, 2024, 19 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202110511245.0, mailed Apr. 23, 2025, 18 pages. [cited by applicant]