IP Library › Granted Patent US 12,744,974
Granted Patent B2
US 12,744,974 · App. 18/656,705 · Granted Sep 22, 2026

Methods and devices for generating customized video segment based on content features

Inventors: Md Ibrahim Khalil (Kanata, CA); Peng Dai (Shenzhen, CN); Hanwen Liang (Kanata, CA); Lizhe Chen (Kanata, CA); Varshanth Ravindra Rao (Kanata, CA); Juwei Lu (Kanata, CA); Songcen Xu (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
H04N21/8549G06V10/70G06V20/46G06V20/49G06F3/0484
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,974
App. No.
18/656,705
Granted
Sep 22, 2026
Kind
B2
Abstract

Methods and devices for generating a customized video segment from a video are disclosed. The video is partitioned into video segments. For each respective video segment, a respective set of scores is computed, where each score represents a respective content feature in the respective video segment. A respective weighted aggregate score is computed for each respective video segment by applying, to each respective set of scores, a common set of weight values. A selected video segment is outputted as the customized video segment, where the selected video segment is selected from one or more high-ranked video segments having high-ranked weighted aggregate scores.

Claims (75)

1 . A method for generating a customized video segment from a video, the method comprising:

computing, for each respective video segment of one or more video segments of the video, each video segment having two or more frames, a respective set of scores, each score representing a respective content feature in the respective video segment;

receiving, from a user device, user input including a selection of one or more user-selected weight values to be applied to each respective set of scores;

computing a respective weighted aggregate score for each respective video segment by applying, to each respective set of scores, a common set of weight values including the user-selected one or more weight values;

selecting a selected video segment as the customized video segment, the selected video segment being selected from one or more high-ranked video segments having high-ranked weighted aggregate scores;

receiving, from the user device, a changed user input that causes the common set of weight values to change to a changed common set of weight values;

recomputing the respective weighted aggregate score for each respective video segment using the changed common set of weight values;

selecting an updated selected video segment as the customized video segment, based on the updated selected video segment being one of one or more reranked video segments having high-ranked recomputed weighted aggregate scores; and

outputting the updated selected video segment as the customized video segment.

2 . The method of claim 1 , wherein the updated selected video segment has a highest ranked recomputed weighted aggregate score.

3 . The method of claim 1 , further comprising:

receiving, from the user device, further user input including a user-submitted query;

comparing a query feature vector representing features of the user-submitted query with a respective video segment feature vector representing features of each respective reranked video segment; and

selecting the reranked video segment represented by the video segment feature vector having a highest similarity with the query feature vector as the updated selected video segment.

4 . The method of claim 1 , further comprising:

prior to receiving the changed user input, providing output to the user device to cause the user device to provide a preview of the selected video segment together with a visual indication of the weighted aggregate score for the selected video segment.

5 . The method of claim 4 , further comprising:

after selecting the updated selected video segment, updating the preview based on the updated selected video segment and updating the visual indication based on the recomputed weighted aggregate score.

6 . The method of claim 1 , wherein outputting the updated selected video segment comprises:

computing a respective frame difference between a start and an end frame of each respective reranked video segment; and

selecting the updated selected video segment to be a reranked video segment having a respective frame difference that falls within a defined difference threshold.

7 . The method of claim 1 , wherein outputting the updated selected video segment comprises:

computing a frame difference between a start and an end frame of the updated selected video segment;

in response to the frame difference exceeding a defined difference threshold, defining a frame previous to the end frame as a new end frame or defining a frame following the start frame as a new start frame; and

repeating the computing and the defining until the frame difference falls within the defined difference threshold.

8 . The method of claim 1 , wherein computing the respective set of scores for each respective video segment comprises, for a given video segment:

generating each respective score in the set of scores by processing the given video segment using a respective trained content feature extraction model.

9 . The method of claim 8 , wherein the respective trained content feature extraction model includes at least one of:

a trained action prediction model;

a trained emotion prediction model;

a trained cheering prediction model;

a trained speed detection model; or

a trained loop detection model.

10 . The method of claim 1 , further comprising:

partitioning the video into the one or more video segments by computing an amount of change between every pair of two consecutive frames of the video; and

defining a start frame of a video segment when the computed amount of change exceeds a defined scene change threshold.

11 . The method of claim 1 , further comprising:

prior to outputting the updated selected video segment as the customized video segment, detecting a region of interest (ROI) in the selected video segment; and

zooming in on the ROI in the updated selected video segment.

12 . The method of claim 1 , further comprising:

prior to outputting the updated selected video segment as the customized video segment, defining a plurality of frames at a start of the updated selected video segment as variable start frames or defining a plurality of frames at an end of the updated selected video segment as variable end frames; and

outputting the customized video segment to have a variable length, wherein the variable length is variable based on a random selection of one of the variable start frames as a first frame of the customized video segment or a random selection of one of the variable end frames as a last frame of the customized video segment.

13 . The method of claim 1 , wherein the customized video segment is outputted in an animated GIF format.

14 . A computing device comprising:

a processing unit configured to execute instructions to cause the computing device to perform a method comprising:

computing, for each respective video segment of one or more video segments of the video, each video segment having two or more frames, a respective set of scores, each score representing a respective content feature in the respective video segment;

receiving, from a user device, user input including a selection of one or more user-selected weight values to be applied to each respective set of scores;

computing a respective weighted aggregate score for each respective video segment by applying, to each respective set of scores, a common set of weight values including the user-selected one or more weight values; selecting a selected video segment as the customized video segment, the selected video segment being selected from one or more high-ranked video segments having high-ranked weighted aggregate scores;

receiving, from the user device, a changed user input that causes the common set of weight values to change to a changed common set of weight values;

recomputing the respective weighted aggregate score for each respective video segment using the changed common set of weight values;

selecting an updated selected video segment as the customized video segment, based on the updated selected video segment being one of one or more reranked video segments having high-ranked recomputed weighted aggregate scores; and

outputting the updated selected video segment as the customized video segment.

15 . The computing device of claim 14 , wherein the instructions cause the computing device to perform the method further comprising:

receiving, from the user device, further user input including a user-submitted query;

comparing a query feature vector representing features of the user-submitted query with a respective video segment feature vector representing features of each respective high-ranked reranked video segment; and

selecting the reranked video segment represented by the video segment feature vector having a highest similarity with the query feature vector as the updated selected video segment.

16 . The computing device of claim 14 , wherein the instructions cause the computing device to perform the method further comprising:

prior to receiving the changed user input, providing output to the computing device to cause the computing device to provide a preview of the selected video segment together with a visual indication of the weighted aggregate score for the selected video segment; and

after selecting the updated selected video segment, updating the preview based on the updated selected video segment and updating the visual indication based on the recomputed weighted aggregate score.

17 . A non-transitory computer readable medium storing instructions thereon, wherein the instructions are executable by a processing unit of a computing device to cause the computing device to perform a method comprising:

computing, for each respective video segment of one or more video segments of the video, each video segment having two or more frames, a respective set of scores, each score representing a respective content feature in the respective video segment;

receiving, from a user device, user input including a selection of one or more user-selected weight values to be applied to each respective set of scores;

computing a respective weighted aggregate score for each respective video segment by applying, to each respective set of scores, a common set of weight values including the user-selected one or more weight values;

selecting a selected video segment as the customized video segment, the selected video segment being selected from one or more high-ranked video segments having high-ranked weighted aggregate scores;

receiving, from the user device, a changed user input that causes the common set of weight values to change to a changed common set of weight values;

recomputing the respective weighted aggregate score for each respective video segment using the changed common set of weight values;

selecting an updated selected video segment as the customized video segment, based on the updated selected video segment being one of one or more reranked video segments having high-ranked recomputed weighted aggregate scores; and

outputting the updated selected video segment as the customized video segment.

18 . The non-transitory computer readable medium of claim 17 , wherein the instructions cause the computing device to perform the method further comprising:

receiving, from the user device, further user input including a user-submitted query;

comparing a query feature vector representing features of the user-submitted query with a respective video segment feature vector representing features of each respective reranked video segment; and

selecting the reranked video segment represented by the video segment feature vector having a highest similarity with the query feature vector as the updated selected video segment.

19 . The non-transitory computer readable medium of claim 17 , wherein the instructions cause the computing device to perform the method further comprising:

prior to receiving the changed user input, providing output to the computing device to cause the computing device to provide a preview of the selected video segment together with a visual indication of the weighted aggregate score for the selected video segment; and

after selecting the updated selected video segment, updating the preview based on the updated selected video segment and updating the visual indication based on the recomputed weighted aggregate score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2024
From: KHALIL, MD IBRAHIM; DAI, PENG; LIANG, HANWEN; CHEN, LIZHE; RAO, VARSHANTH RAVINDRA; LU, JUWEI; XU, SONGCEN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 067331/0088 →
Continuity (2)
Continuation PCTCN2022070589 · Jan 6, 2022
Related Publication 20240292073A1 · Aug 29, 2024
References Cited (22)
US 9646227B2 · Suri · 2017 [cited by examiner]
US 9704231B1 · Kulewski · 2017 [cited by examiner]
US 9782678B2 · Long et al. · 2017 [cited by applicant]
US 11468675B1 · Agarwal · 2022 [cited by examiner]
US 11763564B1 · Sadoughi Nourabadi · 2023 [cited by examiner]
US 20160070962A1 · Shetty · 2016 [cited by examiner]
US 20190297367A1 · Osminer · 2019 [cited by applicant]
US 20210125359A1 · Dolce · 2021 [cited by examiner]
US 20220215560A1 · Ji · 2022 [cited by examiner]
US 20230110002A1 · Luo · 2023 [cited by examiner]
CN 110191358A · 2019 [cited by applicant]
CN 111126262A · 2020 [cited by applicant]
CN 111698554A · 2020 [cited by applicant]
CN 112445935A · 2021 [cited by applicant]
WO 2019293869A1 · 2019 [cited by applicant]
English Translation of Chinese Publication CN112445935 Mar. 2021 (Year: 2021). [cited by examiner]
English Translation of Chinese Publication CN111783731A Oct. 2020 (Year: 2020). [cited by examiner]
Chinese Publication CN111078943 Apr. 2020 (Year: 2020). [cited by examiner]
English translation of Chinese Publication CN113516050 Oct. 2021 (Year: 2021). [cited by examiner]
English translation of Chinese Publication CN111738173 Oct. 2020 (Year: 2020). [cited by examiner]
International Search Report, WO2023/130326 A1, Jul. 13, 2023. [cited by applicant]
Subramanian et al., “Automatically detect sports highlights in video with Amazon SageMaker”, Amazon Machine Learning, Amazon SageMaker, Artificial Intelligence, Sports, Nov. 12, 2021. [cited by applicant]