IP Library Granted Patent US 10,572,735
Granted Patent B2
US 10,572,735 · App. 14/675,464 · Granted Feb 25, 2020

Detect sports video highlights for mobile computing devices

Inventors: Zheng Han (Beijing, CN); Xiaowei Dai (Beijing, CN); Xianjun Huang (Beijing, CN); Fan Yang (Beijing, CN)
Assignee: BEIJING SHUNYUAN KAIHUA TECHNOLOGY LIMITED
G06K9/00724G06K9/00751G06K9/4628G11B27/031G11B27/034G11B27/036G11B27/06G11B27/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,572,735
App. No.
14/675,464
Granted
Feb 25, 2020
Kind
B2
Abstract

A solution is provided for detecting in real time video highlights in a sports video at a mobile computing device. A highlight detection module of the mobile computing device extracts visual features from each video frame of the sports video using a trained feature model and detects a highlight in the video frame based on the extracted visual features of the video frame using a trained detection model. The feature model and detection model are trained with a convolutional neural network on a large corpus of videos to generate category level and pair-wise frame feature vectors. Based on the detection, the highlight detection module generates a highlight score for each video frame of the sports video and presents the highlight scores to users of the computing device. The feature model and detection model are dynamically updated based on the real time highlight detection data collected by the mobile computing device.

Claims (77)

1. A computer-implemented method for detecting highlights in a sports video of a sport at a mobile computing device, comprising:

receiving a sports video having a plurality of video frames at the mobile computing device; and

for each video frame of the plurality of video frames:

extracting, using a feature model that is trained to identify classes of sports in a single video frame, a plurality of visual features of the video frame, the feature model is trained, using images of the sports, to extract frame-based features;

identifying a class of sport of the video frame;

identifying pair-wise frame feature vectors that are for the class of sport of the video frame, each pair-wise frame feature vectors comprising:

a first feature vector describing visual characteristics of a first video frame having a highlight and having the class of sport, and

a second feature vector describing visual characteristics of a second video frame having no highlight and having the class of sport,

wherein the first video frame and the second video frame are images of the class of sport of the video frame, and

wherein the pair-wise frame feature vectors are generated during a training phase and based on a model previously trained using a training set comprising first sports images of the class of sport that include highlights and second sports images of the class of sport that do not include highlights, and first frame-based features extracted by the feature model and corresponding to the first sports images, and second frame-based features extracted by the feature model and corresponding to the second sports images; and

generating a highlight score for the video frame by:

determining first distances between the extracted visual features and respective first feature vectors of the pair-wise frame feature vectors;

determining second distances between the extracted visual features and respective second feature vectors of the pair-wise frame feature vector; and

combining the first distances and the second distances to generate the highlight score for the video frame.

2. The method of claim 1 , wherein the feature model is trained with a convolutional neural network on a large corpus of videos, and wherein the trained feature model is configured to classify the large corpus of videos into a plurality of classes, and each class of the large corpus of videos is associated with a plurality feature vectors describing category level visual characteristics of the class.

3. The method of claim 1 , wherein the highlight score of the video frame represents a prediction of the video frame having a highlight.

4. The method of claim 1 , further comprising:

presenting the highlight scores of the plurality of video frames of the sports video in a graphical user interface; and

monitoring user interactions with the presented highlight scores of the plurality of video frames of the sports video.

5. The method of claim 4 , further comprising:

responsive to detecting a user interaction with a highlight score of a video frame of the sports video, providing highlight detection data of the sports video and user interaction information to a computer server to update a trained detection model.

6. The method of claim 1 , further comprising:

providing highlight detection data of the sports video to a computer server to update a trained feature model.

7. The method of claim 1 , wherein the class of sport of the video frame is identified based on the extracted plurality of visual features of the video frame.

8. A non-transitory computer readable storage medium storing executable computer program instructions for or detecting highlights in a sports video of a sport at a mobile computing device, the instructions when executed by a computer processor cause the computer processor to:

receive a sports video having a plurality of video frames at the mobile computing device; and

for each video frame of the plurality of video frames:

extract, using a feature model that is trained to identify classes of sports in a single video frame, a plurality of visual features of the video frame, the feature model is trained, using images of the sports, to extract frame-based features;

identify a class of sport of the video frame;

identify pair-wise frame feature vectors that are for the class of sport of the video frame, each pair-wise frame feature vectors comprising:

a first feature vector describing visual characteristics of a first video frame having a highlight and having the class of sport, and

a second feature vector describing visual characteristics of a second video frame having no highlight and having the class of sport,

wherein the first video frame and the second video frame are images of the class of sport of the video frame, and

wherein the pair-wise frame feature vectors are generated during a training phase and based on a model previously trained using

 a training set comprising first sports images of the class of sport that include highlights and second sports images of the class of sport that do not include highlights, and

 first frame-based features extracted by the feature model and corresponding to the first sports images, and second frame-based features extracted by the feature model and corresponding to the second sports images; and

generate a highlight score for the video frame by:

determining first distances between the extracted visual features and respective first feature vectors of the pair-wise frame feature vectors;

determining second distances between the extracted visual features and respective second feature vectors of the pair-wise frame feature vector; and

combining the first distances and the second distances to generate the highlight score for the video frame.

9. The computer readable storage medium of claim 8 , wherein the feature model is trained with a convolutional neural network on a large corpus of videos, and wherein the trained feature model is configured to classify the large corpus of videos into a plurality of classes, and each class of the sports video is associated with a plurality feature vectors describing category level visual characteristics of the class.

10. The computer readable storage medium of claim 8 , wherein the highlight score of the video frame represents a prediction of the video frame having a highlight.

11. The computer readable storage medium of claim 8 , further computer program instructions when executed by a computer processor cause the computer processor to:

present the highlight scores of the plurality of video frames of the sports video in a graphical user interface; and

monitor user interactions with the presented highlight scores of the plurality of video frames of the sports video.

12. The computer readable storage medium of claim 11 , further computer program instructions when executed by a computer processor cause the computer processor to:

responsive to detecting a user interaction with a highlight score of a video frame of the sports video, provide highlight detection data of the sports video and user interaction information to a computer server to update a trained detection model.

13. The computer readable storage medium of claim 8 , further computer program instructions when executed by a computer processor cause the computer processor to:

provide highlight detection data of the sports video to a computer server to update a trained feature model.

14. An apparatus for detecting highlights in a sports video of a sport, comprising:

a non-transitory memory; and

a processor configured to execute instructions stored in the non-transitory memory to:

receive a sports video having a plurality of video frames at the apparatus; and

for each video frame of the plurality of video frames:

extract, using a feature model that is trained to identify classes of sports in a single video frame, a plurality of visual features of the video frame, the feature model is trained, using images of the sports, to extract frame-based features;

identify a class of sport of the video frame;

identify pair-wise frame feature vectors that are for the class of sport of the video frame, each pair-wise frame feature vectors comprising:

a first feature vector describing visual characteristics of a first video frame having a highlight and having the class of sport, and

a second feature vector describing visual characteristics of a second video frame having no highlight and having the class of sport,

wherein the first video frame and the second video frame are images of the class of sport of the video frame, and

wherein the pair-wise frame feature vectors are generated during a training phase and based on a model previously trained using

a training set comprising first sports images of the class of sport that include highlights and second sports images of the class of sport that do not include highlights, and

first frame-based features extracted by the feature model and corresponding to the first sports images, and second frame-based features extracted by the feature model and corresponding to the second sports images; and

generate a highlight score for the video frame by:

determining first distances between the extracted visual features and respective first feature vectors of the pair-wise frame feature vectors;

determining second distances between the extracted visual features and respective second feature vectors of the pair-wise frame feature vector; and

combining the first distances and the second distances to generate the highlight score for the video frame.

15. The apparatus of claim 14 , wherein the feature model is trained with a convolutional neural network on a large corpus of videos, and wherein the trained feature model is configured to classify the large corpus of videos into a plurality of classes, and each class of the large corpus of videos is associated with a plurality feature vectors describing category level visual characteristics of the class.

16. The apparatus of claim 14 , wherein the highlight score of the video frame represents a prediction of the video frame having a highlight.

17. The apparatus of claim 14 , wherein the processor is further configured to execute the instructions stored in the non-transitory memory to:

present the highlight scores of the plurality of video frames of the sports video in a graphical user interface; and

monitor user interactions with the presented highlight scores of the plurality of video frames of the sports video.

18. The apparatus of claim 17 , wherein the processor is further configured to execute the instructions stored in the non-transitory memory to:

responsive to detecting a user interaction with a highlight score of a video frame of the sports video, provide highlight detection data of the sports video and user interaction information to a computer server to update a trained detection model.

19. The apparatus of claim 14 , wherein the processor is further configured to execute the instructions stored in the non-transitory memory to:

provide highlight detection data of the sports video to a computer server to update a trained feature model.

20. The apparatus of claim 14 , wherein the class of sport of the video frame is identified based on the extracted plurality of visual features of the video frame.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2018
From: HUAMI HK LIMITED
To: BEIJING SHUNYUAN KAIHUA TECHNOLOGY LIMITED
Reel/Frame 047175/0408 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2018
From: ZEPP LABS, INC.
To: HUAMI HK LIMITED
Reel/Frame 046756/0986 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2015
From: HAN, ZHENG; DAI, XIAOWEI; HUANG, XIANJUN; YANG, FAN
To: ZEPP LABS, INC.
Reel/Frame 035312/0097 →
Continuity (1)
Related Publication 20160292510A1 · Oct 6, 2016