IP Library Patent Application 15600404
Patent Application
App. No. 15/600,404

METHODS AND SYSTEMS OF SPATIOTEMPORAL PATTERN RECOGNITION FOR VIDEO CONTENT DEVELOPMENT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
15/600,404
Abstract

A system for enabling user interaction with video content includes an ingestion facility configured to access at least one video feed and a machine learning system configured to process the at least one video feed through a spatiotemporal pattern recognition algorithm that applies machine learning on an event in the at least one feed in order to develop an understanding of the event including identifying context information relating to the event and an entry in a relationship library at least detailing a relationship between two visible video features. The system further includes an extraction facility configured to automatically extract content displaying the event and associate the extracted content with the context information, and a video production facility configured to produce a video content data structure that includes the context information. The system further includes a user interface configured with video interaction options that are based on the context information.

Claims (52)

1 . A system for enabling user interaction with video content, comprising:

an ingestion facility executing on at least one processor configured to access at least one video feed;

a machine learning system configured to process the at least one video feed through a spatiotemporal pattern recognition algorithm that applies machine learning on an event in the at least one feed in order to develop an understanding of the event within the at least one video feeds, wherein the understanding includes identifying context information relating to the event and an entry in a relationship library at least detailing a relationship between two visible features of the at least one video feed;

an extraction facility configured to automatically, under computer control, extract content displaying the event and associate the extracted content with the context information;

a video production facility configured to produce a video content data structure that includes the context information; and

an application having a user interface configured to permit a user to interact with the video content data structure, wherein the user interface is further configured with options for user interaction that are based on the context information.

2 . The system of claim 1 , wherein the application is a mobile application.

3 . The system of claim 1 , wherein the application is at least one of a smart television application, a virtual reality headset application and an augmented reality application.

4 . The system of claim 1 , wherein the user interface is a touch screen interface.

5 . The system of claim 4 , wherein the user interface is configured to permit a user to enhance the video feed by selecting a content element to be added to the video feed.

6 . The system of claim 5 , wherein the content element is at least one of a metric and a graphic element that is based on the understanding developed with the machine learning.

7 . The system of claim 1 , wherein the user interface is configured to permit the user to select content for a particular player of a sports event.

8 . The system of claim 1 , wherein the user interface is configured to permit the user to select content relating to a context involving a matchup of two particular players in a sports event.

9 . The system of claim 1 , wherein the system takes at least two video feeds from different time periods, the machine learning facility determines a context that includes a similarity between at least one of a plurality of players and a plurality of plays in the two feeds, and wherein the user interface is configured to permit the user to select at least one of the players and the plays to obtain a video feed that illustrates a comparison.

10 . The system of claim 1 , wherein the user interface includes options for at least one of editing, cutting and sharing a video clip that includes the video data structure.

11 . The system of claim 1 , wherein the at least one video feed comprises 3D motion camera data captured from a live sports venue.

12 . The system of claim 1 , wherein the machine learning facility increases its ability to develop the understanding by ingesting a plurality of events for which context has already been identified.

13 . The system of claim 1 , wherein using machine learning to develop the understanding of the event further comprises using events in position tracking data over time obtained from at least one of the at least one video feed and a chip-based player tracking system and wherein the understanding is based on at least two of spatial configuration, relative motion, and projected motion of at least one of a player and an item used in a game.

14 . The system of claim 1 , wherein using the machine learning to develop the understanding of the event further comprises aligning multiple unsynchronized input feeds related to the event using at least one of a hierarchy of algorithms and a hierarchy of human operators, wherein the unsynchronized input feeds are selected from the group consisting of one or more broadcast video feeds of the event, one or more feeds of tracking video for the event, and one or more play-by-play data feeds of the event.

15 . The system of claim 14 , wherein the multiple unsynchronized input feeds include at least three feeds selected from at least two types related to the event.

16 . The system of claim 14 , further comprising at least one of validating and modifying the alignment of the unsynchronized input feeds using a hierarchy involving at least two of one or more algorithms, one or more human operators, and one or more input feeds.

17 . The system of claim 1 , further comprising at least one of validating the understanding and modifying the understanding using a hierarchy involving at least two of one or more algorithms, one or more human operators, and one or more input feeds.

18 . The system of claim 11 , further comprising automatically developing a semantic index of a video feed based on the machine understanding of the event in the video feed to indicate a time of the event in the video feed and a location of a display of the event in the video feed.

19 . The system of claim 18 , wherein the location of the display of the event in the video feed includes at least one of a pixel location, a voxel location, a raster image location.

20 . The system of claim 18 , further comprising providing the semantic index of the video feed with the video feed to enable augmentation of the video feed.

21 . The system of claim 20 , wherein augmentation of the video feed includes adding content based on to the location of the display and enabling at least one of a touch interface feature and a mouse interface feature based on the identified location.

22 . The system of claim 1 , wherein extracting the content displaying the event includes automatically extracting a cut from the at least one video feed using a combination of the understanding developed with the machine learning and an understanding developed with the machine learning of another input feed selected from the group consisting of a broadcast video feed, an audio feed, and a closed caption feed.

23 . The system of claim 22 , wherein the understanding developed with machine learning of the other input feed includes at least one of a portion of content of a broadcast commentary and a change in camera view in the input feed.

24 . A method for enabling a mobile application allowing user interaction with video content, comprising:

taking at least one video feed;

processing the at least one video feed through a spatiotemporal pattern recognition algorithm that uses machine learning to develop an understanding of an event within the at least one video feed, wherein the understanding includes identifying context information relating to the event and an entry in a relationship library at least detailing a relationship between two visible features of the at least one video feed;

automatically, under computer control, extracting content displaying the event and associating the extracted content with the context information;

producing a video content data structure that includes the context information; and

providing a mobile application having a user interface configured to permit a user to interact with the video content data structure, wherein the user interface is configured to include options for user interaction based on the context information.

25 . The method of claim 24 , wherein the user interface is a touch screen interface.

26 . The method of claim 25 , wherein the user interface is configured to permit a user to enhance the video feed by selecting a content element to be added to the video feed.

27 . The method of claim 26 , wherein the content element is at least one of a metric and a graphic element that is based on the machine understanding.

28 . The method of claim 24 , wherein the user interface is configured to permit the user to select content for a particular player of a sports event.

29 . The method of claim 24 , wherein the user interface is configured to permit the user to select content relating to a context involving the matchup of two particular players in a sports event.

30 . The method of claim 24 , further comprising taking at least two video feeds from different time periods, the machine learning facility determines a context the includes a similarity between at least one of a plurality of players and a plurality of plays in the at least two feeds and the user interface is configured to permit the user to select at least one of the players and the plays to obtain a video feed that illustrates a comparison.

31 . The method of claim 24 , wherein the user interface includes options for at least one of editing, cutting and sharing a video clip that includes the video data structure.

32 . The method of claim 24 , wherein the video feed comprises 3D motion camera data captured from a live sports venue.

33 . The method of claim 24 , wherein the machine learning facility increases its ability to develop the understanding by ingesting a plurality of events for which context has already been identified.

34 . The method of claim 24 , wherein using the machine learning to develop the understanding of the event further comprises using events in position tracking data over time obtained from at least one of the at least one video feed and a chip-based player tracking system, and wherein the understanding is based on at least two of spatial configuration, relative motion, and projected motion of at least one of a player and an item used in a game.

35 . The method of claim 24 , wherein using the machine learning to develop the understanding of the event further comprises aligning multiple unsynchronized input feeds related to the event using at least one of a hierarchy of algorithms and a hierarchy of human operators, wherein the unsynchronized input feeds are selected from the group consisting of one or more broadcast video feeds of the event, one or more feeds of tracking video for the event, and one or more play-by-play data feeds of the event.

36 . The method of claim 35 , wherein the multiple unsynchronized input feeds include at least three feeds selected from at least two types related to the event.

37 . The method of claim 35 , further comprising at least one of validating and modifying the alignment of the unsynchronized input feeds using a hierarchy involving at least two of one or more algorithms, one or more human operators, and one or more input feeds.

38 . The method of claim 24 , further comprising at least one of validating and modifying the understanding using a hierarchy involving at least two of one or more algorithms, one or more human operators, and one or more input feeds.

39 . The method of claim 24 , further comprising automatically developing a semantic index of a video feed based on the understanding developed with the machine learning of at least one event in the video feed to indicate a time of the event in the video feed and a location of a display of the event in the video feed.

40 . The method of claim 39 , wherein the location of the display of the event in the video feed includes at least one of a pixel location, a voxel location, a raster image location.

41 . The method of claim 39 , further comprising providing the semantic index of the video feed with the video feed to enable augmentation of the video feed.

42 . The method of claim 41 , wherein augmentation of the video feed includes adding content based on to the location of the display and enabling at least one of a touch interface feature and a mouse interface feature based on the identified location.

Assignments (2)
MERGER Recorded Sep 17, 2021
From: SECOND SPECTRUM, INC.
To: GENIUS SPORTS SS, LLC
Reel/Frame 057509/0582 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2017
From: CHANG, YU-HAN; MAHESWARAN, RAJIV; SU, JEFFREY WAYNE; HOLLINGSWORTH, NOEL
To: SECOND SPECTRUM, INC.
Reel/Frame 042505/0633 →